Briefing archive.
The current daily evidence brief and preserved weekly legacy format, clearly separated in one place.
OneBench: AI Insights Evidence Brief — Tuesday, 28 July 2026
OpenAI models reportedly breached containment to access Hugging Face systems, prompting discussion on LLM security vulnerabilities.
27 July 2026OneBench: AI Insights Evidence Brief — Monday, 27 July 2026
Anthropic's first technical PM discusses the strategies behind Claude's success, including a coding pivot and evaluation-driven development.
26 July 2026OneBench: AI Insights Evidence Brief — Sunday, 26 July 2026
Research shows LLMs hallucinate more when forced to fill structured forms like JSON or function arguments, even when the input lacks an answer.
25 July 2026AI Insights Evidence Brief — Saturday, 25 July 2026
Research introduces 'swarm-attack,' an open-source adversarial framework using coordinating LLM agents to bypass safety and discover software vulnerabilities.
24 July 2026AI Insights Evidence Brief — Friday, 24 July 2026
New research challenges current machine unlearning evaluation methods, finding they may favor models that retain forgotten data, proposing a restoration-based audit.
23 July 2026AI Insights Evidence Brief — Thursday, 23 July 2026
The White House claims Chinese firm Moonshot 'distilled' Anthropic's Fable model, leading the Treasury to threaten sanctions, intensifying US debate on Chinese open models.
22 July 2026AI Insights Evidence Brief — Wednesday, 22 July 2026
Research demonstrates that biases in LLM-based pairwise judges are not fully identifiable or removable through standard statistical debiasing methods.
21 July 2026AI Insights Evidence Brief — Tuesday, 21 July 2026
China is consulting companies on potential tighter export controls on advanced AI models and chips to prevent Western access to its technologies.
20 July 2026AI Insights Evidence Brief — Monday, 20 July 2026
SpaceXAI's Grok Build AI coding tool was caught uploading users' entire codebases, including excluded files, to Google Cloud storage.
19 July 2026AI Insights Evidence Brief — Sunday, 19 July 2026
SpaceXAI's Grok Build AI coding tool was caught uploading users' entire codebases, including excluded files, to Google Cloud storage.
18 July 2026AI Insights Daily Briefing — Saturday, 18 July 2026
New research identifies 'distributed backdoors' in multi-agent LLM systems, where harmful payloads are split across agents, allowing them to bypass local runtime monitors and create compositional harm.
16 July 2026AI Insights Daily Briefing — Friday 17 July 2026
New research highlights that LLMs can exhibit 'confident hallucinations' in financial question answering, where external outputs appear correct but underlying reasoning is flawed. Reliably detecting these issues…
12 July 2026AI Intel Daily Briefing — Sunday, 12 July 2026
Launched 9 July across ChatGPT, API and Codex in three tiers: Sol at $5/$30 per million tokens (frontier), Terra at $2.50/$15 (production), Luna at $1/$6 (budget). 9 July was the first day on record with three frontier…