This frozen page shows Augur's claims and source links for one sent dispatch. Stored spot-checks appear only where the frozen edition supports them; absence is not presented as verification.
Epoch AI My view is more urgent: today's evidence shows that old weaknesses become dangerous when agents erase the labor cost of exploiting them.
Assertion status: No spot-check verdict is published for this assertion.
Multiple independent benchmarks, including ExploitGym, ExploitBench, UK AI Security Institute's Cyber Ranges, Irregular's CyScenarioBench and FrontierCyber, have shown frontier AI models can discover security vulnerabilities in real-world code and develop exploits to hack realistic systems.
No stored spot-check names this claim in this edition.
Assertion 2
Princeton CITP The model did not invent the flaw.
Assertion status: No spot-check verdict is published for this assertion.
A 2022 research report identified a critical privacy vulnerability in certain ballot scanning machines where the algorithm used to assign random numbers to electronic ballot records was deterministic and reversible, allowing the sequence of ballot casting to be reconstructed.
No stored spot-check names this claim in this edition.
Author Max Springer used AI coding agents and public records to reconstruct the casting order of 1.52 million in-person ballots (98.9% of the total) from Georgia's May 2026 primary election.
No stored spot-check names this claim in this edition.
In 114 of the 139 Georgia counties examined, the in-person casting order of ballots could be recovered from public files alone, enabling the identification of voters' ballots in specific contexts.
No stored spot-check names this claim in this edition.
Assertion 3
Import AI That is not an outbreak.
Assertion status: No spot-check verdict is published for this assertion.
Researchers from the University of Toronto, the Vector Institute, the University of Cambridge, and ServiceNow built a prototype computer virus that uses AI models to compromise computers and leverages their GPU resources for inference to infect more hosts.
No stored spot-check names this claim in this edition.
The proof-of-concept worm operates using an open-weight LLM published in 2025 that fits on a single A100 GPU with 80GB of VRAM, without relying on monitored vendor APIs.
No stored spot-check names this claim in this edition.
The prototype worm achieved approximately 80% success in vulnerability detection, 53% in exploitation, and 88% in self-replication, resulting in an overall attack success rate of about 37%.
No stored spot-check names this claim in this edition.
Assertion 4
Embrace The Red An IT leader should treat routing credentials as production secrets and keep agent permissions narrow even when the underlying model is trusted.
Assertion status: No spot-check verdict is published for this assertion.
A malicious actor can hijack LiteLLM traffic, steal backend LLM provider keys, and inject tool calls by exploiting legitimate proxy-admin credentials and routing configuration endpoints.
No stored spot-check names this claim in this edition.
Assertion 5
Microsoft Research OpenRouter separately launched Ori Eval to compare models on a builder's own prompts and data.
Assertion status: No spot-check verdict is published for this assertion.
Microsoft Research introduced Orchard, an open-source framework centered on Orchard Env, a Kubernetes-based environment service designed to support scalable agentic AI research across multiple task domains.
No stored spot-check names this claim in this edition.
Orchard-SWE achieved a 69.7% score on the SWE-bench Verified benchmark using approximately 3 billion active parameters, reaching 73.0% with value-model reranking.
No stored spot-check names this claim in this edition.
Orchard-GUI, a 4-billion-parameter vision-language model, achieved an average success rate of 68.4% across the WebVoyager, Online-Mind2Web, and DeepShop benchmarks.
No stored spot-check names this claim in this edition.
Assertion 6
OpenRouter Orchard supplies exactly the kind of headline benchmark number an engineer should re-run on their own tasks, and Ori Eval is a tool for doing that comparison.
Assertion status: No spot-check verdict is published for this assertion.
OpenRouter launched Ori Eval, a tool designed to help users identify the best AI model for their specific application by running evaluations on the user's own prompts and data.
No stored spot-check names this claim in this edition.
Assertion 7
ChinaAI The capability pressure has been building since mid-July, and a newer comparison put DeepSeek R1 only a few points behind leading closed models, displacing an earlier concern that it could not keep up.
Assertion status: No spot-check verdict is published for this assertion.
Deploying Moonshot AI's K3 model requires at least 16 H200 GPUs, a hardware cost that makes personal deployment unaffordable for most users.
No stored spot-check names this claim in this edition.
Running 64 K3 accelerator cards at full load consumes 45 kilowatts of electricity, exceeding the capacity of ordinary household electrical meters and wiring.
No stored spot-check names this claim in this edition.
The gap between open-weight model intelligence and proprietary model intelligence has narrowed significantly, with Deepseek R1 being only a few points behind leading models.
No stored spot-check names this claim in this edition.
Assertion 10
- Georgia election officials should publish any scanner or records-format changes that prevent casting-order reconstruction before the next public export. Princeton CITP
Assertion status: No spot-check verdict is published for this assertion.
A 2022 research report identified a critical privacy vulnerability in certain ballot scanning machines where the algorithm used to assign random numbers to electronic ballot records was deterministic and reversible, allowing the sequence of ballot casting to be reconstructed.
No stored spot-check names this claim in this edition.
Author Max Springer used AI coding agents and public records to reconstruct the casting order of 1.52 million in-person ballots (98.9% of the total) from Georgia's May 2026 primary election.
No stored spot-check names this claim in this edition.
In 114 of the 139 Georgia counties examined, the in-person casting order of ballots could be recovered from public files alone, enabling the identification of voters' ballots in specific contexts.
No stored spot-check names this claim in this edition.
Assertion 11
- LiteLLM maintainers should add public guidance on admin-route controls, provider-key isolation, and tool-call logging. Embrace The Red
Assertion status: No spot-check verdict is published for this assertion.
A malicious actor can hijack LiteLLM traffic, steal backend LLM provider keys, and inject tool calls by exploiting legitimate proxy-admin credentials and routing configuration endpoints.
No stored spot-check names this claim in this edition.
Assertion 12
- Independent teams need to reproduce Orchard's coding and browser results, including the cost of ranking multiple answers. Microsoft Research
Assertion status: No spot-check verdict is published for this assertion.
Microsoft Research introduced Orchard, an open-source framework centered on Orchard Env, a Kubernetes-based environment service designed to support scalable agentic AI research across multiple task domains.
No stored spot-check names this claim in this edition.
Orchard-SWE achieved a 69.7% score on the SWE-bench Verified benchmark using approximately 3 billion active parameters, reaching 73.0% with value-model reranking.
No stored spot-check names this claim in this edition.
Orchard-GUI, a 4-billion-parameter vision-language model, achieved an average success rate of 68.4% across the WebVoyager, Online-Mind2Web, and DeepShop benchmarks.
No stored spot-check names this claim in this edition.
Running 64 K3 accelerator cards at full load consumes 45 kilowatts of electricity, exceeding the capacity of ordinary household electrical meters and wiring.
No stored spot-check names this claim in this edition.
Running 64 K3 accelerator cards at full load consumes 45 kilowatts of electricity, exceeding the capacity of ordinary household electrical meters and wiring.
No stored spot-check names this claim in this edition.
Researchers from the University of Toronto, the Vector Institute, the University of Cambridge, and ServiceNow built a prototype computer virus that uses AI models to compromise computers and leverages their GPU resources for inference to infect more hosts.
No stored spot-check names this claim in this edition.
The proof-of-concept worm operates using an open-weight LLM published in 2025 that fits on a single A100 GPU with 80GB of VRAM, without relying on monitored vendor APIs.
No stored spot-check names this claim in this edition.
The prototype worm achieved approximately 80% success in vulnerability detection, 53% in exploitation, and 88% in self-replication, resulting in an overall attack success rate of about 37%.
No stored spot-check names this claim in this edition.
Microsoft Research introduced Orchard, an open-source framework centered on Orchard Env, a Kubernetes-based environment service designed to support scalable agentic AI research across multiple task domains.
No stored spot-check names this claim in this edition.
Orchard-SWE achieved a 69.7% score on the SWE-bench Verified benchmark using approximately 3 billion active parameters, reaching 73.0% with value-model reranking.
No stored spot-check names this claim in this edition.
Orchard-GUI, a 4-billion-parameter vision-language model, achieved an average success rate of 68.4% across the WebVoyager, Online-Mind2Web, and DeepShop benchmarks.
No stored spot-check names this claim in this edition.
A 2022 research report identified a critical privacy vulnerability in certain ballot scanning machines where the algorithm used to assign random numbers to electronic ballot records was deterministic and reversible, allowing the sequence of ballot casting to be reconstructed.
No stored spot-check names this claim in this edition.
Author Max Springer used AI coding agents and public records to reconstruct the casting order of 1.52 million in-person ballots (98.9% of the total) from Georgia's May 2026 primary election.
No stored spot-check names this claim in this edition.
In 114 of the 139 Georgia counties examined, the in-person casting order of ballots could be recovered from public files alone, enabling the identification of voters' ballots in specific contexts.