Augur Dispatch

Chain of evidence

Evidence for 2026-08-13

This frozen page shows Augur's claims and source links for one sent dispatch. Stored spot-checks appear only where the frozen edition supports them; absence is not presented as verification.

As of:

Bundle identity: evidence-bundle-v1-9dc6564936a0f45d5a534e85d2c71de52b359a85834f1d9977f840b5c928229e

Format: evidence-bundle-v1 · 20 claims

Assertion 1

Switching on thinking, the mode where a model reasons step by step before answering, recovered roughly 40 to 65 percent of the stored-but-unretrieved facts in models tuned for it, but 11 to 12 percent still refused to surface even then, and the 2 to 5 percent of facts that were never encoded stay out of reach no matter how hard the model thinks Google Research Blog.

Assertion status: No spot-check verdict is published for this assertion.

The authors state that Gemini-3-Pro and GPT-5 encode 95–98% of facts but fail to directly recall 26–34% of them.

Claim 45724 Label: fact Provenance: primary Recorded

Google Research Blog

No stored spot-check names this claim in this edition.

The authors claim that even with 'thinking' enabled, Gemini-3-Pro and GPT-5 fail to recall 11–12% of facts.

Claim 45725 Label: fact Provenance: primary Recorded

Google Research Blog

No stored spot-check names this claim in this edition.

The authors state that 'thinking' recovers roughly 40–65% of encoded-but-not-directly-known facts in thinking-optimized models.

Claim 45728 Label: fact Provenance: primary Recorded

Google Research Blog

No stored spot-check names this claim in this edition.

Assertion 2

It extends earlier Google work showing models like Gemini-2.5 could recall answers that were nearly unrecoverable with reasoning switched off Google Research Blog.

Assertion status: No spot-check verdict is published for this assertion.

Models such as Gemini-2.5 and Qwen3-32B successfully recall answers that are virtually unrecoverable when their reasoning capabilities are disabled.

Claim 26892 Label: fact Provenance: primary Recorded

Google Research Blog

No stored spot-check names this claim in this edition.

Assertion 3

It also fits the wider record: models have long scored better on benchmarks than in everyday use, and a Princeton study concluded a whole generation of flagships was not meaningfully more reliable than its predecessors, contradicting the vendors' launch claims Nate Jones Latent Space.

Assertion status: No spot-check verdict is published for this assertion.

Google released Gemma 4 Quantization-Aware Training (QAT) checkpoints for local deployment.

Claim 1316 Label: fact Provenance: primary Recorded

Latent Space

No stored spot-check names this claim in this edition.

A Princeton ICML 2026 paper update concludes that GPT 5.5, Gemini 3.1 Pro/3.5 Flash, and Claude Opus 4.7 are not meaningfully more reliable than previous models.

Claim 1318 Label: fact Provenance: primary Recorded

Latent Space

No stored spot-check names this claim in this edition.

Current large language models perform significantly better on benchmarks than in practical, everyday scenarios.

Claim 13507 Label: fact Provenance: primary Recorded

Nate Jones

No stored spot-check names this claim in this edition.

Assertion 4

More than 180 organizations have signed the Code of Practice that turns those duties into concrete steps EU AI Act Newsletter.

Assertion status: No spot-check verdict is published for this assertion.

The European Commission's AI Office, alongside national authorities, began enforcing the EU AI Act and its new transparency requirements starting on 2 August 2026.

Claim 45701 Label: fact Provenance: primary Recorded

EU AI Act Newsletter

No stored spot-check names this claim in this edition.

Under the new transparency rules taking effect on 2 August 2026, chatbots and interactive systems must disclose to users that they are interacting with AI, deepfakes must be labeled, and AI-generated or altered content must carry machine-readable marks for detection.

Claim 45702 Label: fact Provenance: primary Recorded

EU AI Act Newsletter

No stored spot-check names this claim in this edition.

The European Commission published a list of more than 180 organizations that have signed the Code of Practice operationalizing the AI Act's transparency rules.

Claim 45703 Label: fact Provenance: primary Recorded

EU AI Act Newsletter

No stored spot-check names this claim in this edition.

Assertion 5

RingCentral joined the customer roster, saying it uses ChatGPT Work and Codex, OpenAI's coding agent, to speed AI product development and pool operating knowledge across engineering and operations teams OpenAI News.

Assertion status: No spot-check verdict is published for this assertion.

RingCentral utilizes ChatGPT Work and Codex to accelerate its AI product development.

Claim 45679 Label: fact Provenance: primary Recorded

OpenAI News

No stored spot-check names this claim in this edition.

RingCentral uses ChatGPT Work and Codex to centralize operational intelligence across its engineering and operations teams.

Claim 45680 Label: fact Provenance: primary Recorded

OpenAI News

No stored spot-check names this claim in this edition.

Assertion 6

OpenAI also published its own research arguing enterprises are moving from AI assistance to AI execution OpenAI News.

Assertion status: No spot-check verdict is published for this assertion.

OpenAI research indicates that enterprises are adopting agentic AI, utilizing ChatGPT and Codex.

Claim 45699 Label: fact Provenance: primary Recorded

OpenAI News

No stored spot-check names this claim in this edition.

Assertion 7

Cursor released Grok 4.5, its coding-and-agents model developed jointly with SpaceXAI, on July 8; five weeks later the pair have shipped Grok 4.6, which Cursor says is meant for agents that run longer and for work with heavier visual and interactive elements, and the cadence is the real news Cursor Blog Latent Space Cursor Blog.

Assertion status: No spot-check verdict is published for this assertion.

xAI launched Grok 4.5, a model positioned for coding and agent workflows, on July 8, 2026.

Claim 10623 Label: fact Provenance: primary Recorded

Latent Space

No stored spot-check names this claim in this edition.

xAI and Cursor partnered to train Grok 4.5.

Claim 10633 Label: fact Provenance: primary Recorded

Latent Space

No stored spot-check names this claim in this edition.

Cursor has released Grok 4.5, developed jointly with SpaceXAI.

Claim 27373 Label: fact Provenance: primary Recorded

Cursor Blog

No stored spot-check names this claim in this edition.

Grok 4.5 is characterized as Cursor's most intelligent model to date.

Claim 27374 Label: opinion Provenance: primary Recorded

Cursor Blog

No stored spot-check names this claim in this edition.

Grok 4.5 is available on Cursor desktop, web, iOS, CLI, and SDK.

Claim 27380 Label: fact Provenance: primary Recorded

Cursor Blog

No stored spot-check names this claim in this edition.

Cursor and SpaceXAI released the Grok 4.6 model on August 12, 2026.

Claim 45659 Label: fact Provenance: primary Recorded

Cursor Blog

No stored spot-check names this claim in this edition.

Assertion 8

- Google Research's encode-versus-recall gap invites replication; the same measurement run on open-weight models, the kind anyone can download and run, would show whether the finding holds beyond these two flagships. Google Research Blog

Assertion status: No spot-check verdict is published for this assertion.

The authors state that Gemini-3-Pro and GPT-5 encode 95–98% of facts but fail to directly recall 26–34% of them.

Claim 45724 Label: fact Provenance: primary Recorded

Google Research Blog

No stored spot-check names this claim in this edition.

The authors claim that even with 'thinking' enabled, Gemini-3-Pro and GPT-5 fail to recall 11–12% of facts.

Claim 45725 Label: fact Provenance: primary Recorded

Google Research Blog

No stored spot-check names this claim in this edition.

The authors state that 'thinking' recovers roughly 40–65% of encoded-but-not-directly-known facts in thinking-optimized models.

Claim 45728 Label: fact Provenance: primary Recorded

Google Research Blog

No stored spot-check names this claim in this edition.

Assertion 9

- The European Commission's signatory list for the transparency Code of Practice, now past 180 organizations, is the roster to scan for holdout labs. EU AI Act Newsletter

Assertion status: No spot-check verdict is published for this assertion.

The European Commission's AI Office, alongside national authorities, began enforcing the EU AI Act and its new transparency requirements starting on 2 August 2026.

Claim 45701 Label: fact Provenance: primary Recorded

EU AI Act Newsletter

No stored spot-check names this claim in this edition.

Under the new transparency rules taking effect on 2 August 2026, chatbots and interactive systems must disclose to users that they are interacting with AI, deepfakes must be labeled, and AI-generated or altered content must carry machine-readable marks for detection.

Claim 45702 Label: fact Provenance: primary Recorded

EU AI Act Newsletter

No stored spot-check names this claim in this edition.

The European Commission published a list of more than 180 organizations that have signed the Code of Practice operationalizing the AI Act's transparency rules.

Claim 45703 Label: fact Provenance: primary Recorded

EU AI Act Newsletter

No stored spot-check names this claim in this edition.

Assertion 10

- Cursor's claim that Grok 4.6 suits long-running agents needs third-party agent evaluations before anyone standardizes on it. Cursor Blog

Assertion status: No spot-check verdict is published for this assertion.

Cursor and SpaceXAI released the Grok 4.6 model on August 12, 2026.

Claim 45659 Label: fact Provenance: primary Recorded

Cursor Blog

No stored spot-check names this claim in this edition.

Assertion 11

- Frontier labs now face evidence that encrypted reasoning traces from their APIs can be decoded and replayed to weaker models; their response will show how seriously they treat leakage of a model's hidden reasoning steps. Latent Space

Assertion status: No spot-check verdict is published for this assertion.

A research paper demonstrates that encrypted reasoning traces from frontier API models can be decoded and replayed to weaker models to extract the hidden chain-of-thought content.

Claim 45719 Label: fact Provenance: primary Recorded

Latent Space

No stored spot-check names this claim in this edition.