Augur Dispatch

Chain of evidence

Evidence for 2026-08-04

This frozen page shows Augur's claims and source links for one sent dispatch. Stored spot-checks appear only where the frozen edition supports them; absence is not presented as verification.

As of:

Bundle identity: evidence-bundle-v1-c532c566fd46552e8fab86ed8ea7837f5596ba84666a63c7d575d775daecd07b

Format: evidence-bundle-v1 · 23 claims

Assertion 1

Epoch AI My view is more urgent: today's evidence shows that old weaknesses become dangerous when agents erase the labor cost of exploiting them.

Assertion status: No spot-check verdict is published for this assertion.

Multiple independent benchmarks, including ExploitGym, ExploitBench, UK AI Security Institute's Cyber Ranges, Irregular's CyScenarioBench and FrontierCyber, have shown frontier AI models can discover security vulnerabilities in real-world code and develop exploits to hack realistic systems.

Claim 35136 Label: fact Provenance: primary Recorded

Epoch AI Gradient Updates

No stored spot-check names this claim in this edition.

Assertion 2

Princeton CITP The model did not invent the flaw.

Assertion status: No spot-check verdict is published for this assertion.

A 2022 research report identified a critical privacy vulnerability in certain ballot scanning machines where the algorithm used to assign random numbers to electronic ballot records was deterministic and reversible, allowing the sequence of ballot casting to be reconstructed.

Claim 41646 Label: fact Provenance: primary Recorded

Princeton CITP - Freedom to Tinker

No stored spot-check names this claim in this edition.

Author Max Springer used AI coding agents and public records to reconstruct the casting order of 1.52 million in-person ballots (98.9% of the total) from Georgia's May 2026 primary election.

Claim 41647 Label: fact Provenance: primary Recorded

Princeton CITP - Freedom to Tinker

No stored spot-check names this claim in this edition.

In 114 of the 139 Georgia counties examined, the in-person casting order of ballots could be recovered from public files alone, enabling the identification of voters' ballots in specific contexts.

Claim 41648 Label: fact Provenance: primary Recorded

Princeton CITP - Freedom to Tinker

No stored spot-check names this claim in this edition.

Assertion 3

Import AI That is not an outbreak.

Assertion status: No spot-check verdict is published for this assertion.

Researchers from the University of Toronto, the Vector Institute, the University of Cambridge, and ServiceNow built a prototype computer virus that uses AI models to compromise computers and leverages their GPU resources for inference to infect more hosts.

Claim 41548 Label: fact Provenance: primary Recorded

Import AI

No stored spot-check names this claim in this edition.

The proof-of-concept worm operates using an open-weight LLM published in 2025 that fits on a single A100 GPU with 80GB of VRAM, without relying on monitored vendor APIs.

Claim 41550 Label: fact Provenance: primary Recorded

Import AI

No stored spot-check names this claim in this edition.

The prototype worm achieved approximately 80% success in vulnerability detection, 53% in exploitation, and 88% in self-replication, resulting in an overall attack success rate of about 37%.

Claim 41551 Label: fact Provenance: primary Recorded

Import AI

No stored spot-check names this claim in this edition.

Assertion 4

Embrace The Red An IT leader should treat routing credentials as production secrets and keep agent permissions narrow even when the underlying model is trusted.

Assertion status: No spot-check verdict is published for this assertion.

A malicious actor can hijack LiteLLM traffic, steal backend LLM provider keys, and inject tool calls by exploiting legitimate proxy-admin credentials and routing configuration endpoints.

Claim 41603 Label: fact Provenance: primary Recorded

Embrace The Red

No stored spot-check names this claim in this edition.

In March 2026, the PyPI package for LiteLLM was compromised and distributed a credential-stealing payload.

Claim 41604 Label: fact Provenance: primary Recorded

Embrace The Red

No stored spot-check names this claim in this edition.

Assertion 5

Microsoft Research OpenRouter separately launched Ori Eval to compare models on a builder's own prompts and data.

Assertion status: No spot-check verdict is published for this assertion.

Microsoft Research introduced Orchard, an open-source framework centered on Orchard Env, a Kubernetes-based environment service designed to support scalable agentic AI research across multiple task domains.

Claim 41611 Label: fact Provenance: primary Recorded

Microsoft Research Blog

No stored spot-check names this claim in this edition.

Orchard-SWE achieved a 69.7% score on the SWE-bench Verified benchmark using approximately 3 billion active parameters, reaching 73.0% with value-model reranking.

Claim 41612 Label: fact Provenance: primary Recorded

Microsoft Research Blog

No stored spot-check names this claim in this edition.

Orchard-GUI, a 4-billion-parameter vision-language model, achieved an average success rate of 68.4% across the WebVoyager, Online-Mind2Web, and DeepShop benchmarks.

Claim 41614 Label: fact Provenance: primary Recorded

Microsoft Research Blog

No stored spot-check names this claim in this edition.

Assertion 6

OpenRouter Orchard supplies exactly the kind of headline benchmark number an engineer should re-run on their own tasks, and Ori Eval is a tool for doing that comparison.

Assertion status: No spot-check verdict is published for this assertion.

OpenRouter launched Ori Eval, a tool designed to help users identify the best AI model for their specific application by running evaluations on the user's own prompts and data.

Claim 41493 Label: fact Provenance: primary Recorded

OpenRouter Blog

No stored spot-check names this claim in this edition.

Assertion 7

ChinaAI The capability pressure has been building since mid-July, and a newer comparison put DeepSeek R1 only a few points behind leading closed models, displacing an earlier concern that it could not keep up.

Assertion status: No spot-check verdict is published for this assertion.

Deploying Moonshot AI's K3 model requires at least 16 H200 GPUs, a hardware cost that makes personal deployment unaffordable for most users.

Claim 41516 Label: fact Provenance: primary Recorded

ChinaAI

No stored spot-check names this claim in this edition.

Moonshot AI recommends deploying K3 on a super-node of at least 64 accelerator cards, with total costs of approximately 17 million RMB.

Claim 41517 Label: fact Provenance: primary Recorded

ChinaAI

No stored spot-check names this claim in this edition.

Running 64 K3 accelerator cards at full load consumes 45 kilowatts of electricity, exceeding the capacity of ordinary household electrical meters and wiring.

Claim 41518 Label: fact Provenance: primary Recorded

ChinaAI

No stored spot-check names this claim in this edition.

Assertion 8

AI Engineer One Useful Thing A narrower capability gap does not make the hardware burden disappear.

Assertion status: No spot-check verdict is published for this assertion.

DeepSeek r1 is missing features compared to other major AI companies and may not keep up in the long term.

Claim 471 Label: forecast Provenance: primary Recorded

One Useful Thing

No stored spot-check names this claim in this edition.

OpenAI's o3 model leads in intelligence, followed closely by o4 mini with high reasoning mode and Deepseek R1.

Claim 15736 Label: fact Provenance: primary Recorded

AI Engineer

No stored spot-check names this claim in this edition.

The gap between open-weight model intelligence and proprietary model intelligence has narrowed significantly, with Deepseek R1 being only a few points behind leading models.

Claim 15741 Label: fact Provenance: primary Recorded

AI Engineer

No stored spot-check names this claim in this edition.

Deepseek and Alibaba are the leading providers of open-weight reasoning and non-reasoning models.

Claim 15742 Label: fact Provenance: primary Recorded

AI Engineer

No stored spot-check names this claim in this edition.

Assertion 9

OpenAI A business leader should not treat that as a transferable return.

Assertion status: No spot-check verdict is published for this assertion.

Circles utilizes the OpenAI API and Codex to deliver AI-native telecommunications experiences.

Claim 41489 Label: fact Provenance: primary Recorded

OpenAI News

No stored spot-check names this claim in this edition.

The implementation of OpenAI technology by Circles resulted in a 22% increase in average revenue per user (ARPU).

Claim 41490 Label: fact Provenance: primary Recorded

OpenAI News

No stored spot-check names this claim in this edition.

The implementation of OpenAI technology by Circles resulted in a 9% reduction in customer churn.

Claim 41491 Label: fact Provenance: primary Recorded

OpenAI News

No stored spot-check names this claim in this edition.

Assertion 10

- Georgia election officials should publish any scanner or records-format changes that prevent casting-order reconstruction before the next public export. Princeton CITP

Assertion status: No spot-check verdict is published for this assertion.

A 2022 research report identified a critical privacy vulnerability in certain ballot scanning machines where the algorithm used to assign random numbers to electronic ballot records was deterministic and reversible, allowing the sequence of ballot casting to be reconstructed.

Claim 41646 Label: fact Provenance: primary Recorded

Princeton CITP - Freedom to Tinker

No stored spot-check names this claim in this edition.

Author Max Springer used AI coding agents and public records to reconstruct the casting order of 1.52 million in-person ballots (98.9% of the total) from Georgia's May 2026 primary election.

Claim 41647 Label: fact Provenance: primary Recorded

Princeton CITP - Freedom to Tinker

No stored spot-check names this claim in this edition.

In 114 of the 139 Georgia counties examined, the in-person casting order of ballots could be recovered from public files alone, enabling the identification of voters' ballots in specific contexts.

Claim 41648 Label: fact Provenance: primary Recorded

Princeton CITP - Freedom to Tinker

No stored spot-check names this claim in this edition.

Assertion 11

- LiteLLM maintainers should add public guidance on admin-route controls, provider-key isolation, and tool-call logging. Embrace The Red

Assertion status: No spot-check verdict is published for this assertion.

A malicious actor can hijack LiteLLM traffic, steal backend LLM provider keys, and inject tool calls by exploiting legitimate proxy-admin credentials and routing configuration endpoints.

Claim 41603 Label: fact Provenance: primary Recorded

Embrace The Red

No stored spot-check names this claim in this edition.

In March 2026, the PyPI package for LiteLLM was compromised and distributed a credential-stealing payload.

Claim 41604 Label: fact Provenance: primary Recorded

Embrace The Red

No stored spot-check names this claim in this edition.

Assertion 12

- Independent teams need to reproduce Orchard's coding and browser results, including the cost of ranking multiple answers. Microsoft Research

Assertion status: No spot-check verdict is published for this assertion.

Microsoft Research introduced Orchard, an open-source framework centered on Orchard Env, a Kubernetes-based environment service designed to support scalable agentic AI research across multiple task domains.

Claim 41611 Label: fact Provenance: primary Recorded

Microsoft Research Blog

No stored spot-check names this claim in this edition.

Orchard-SWE achieved a 69.7% score on the SWE-bench Verified benchmark using approximately 3 billion active parameters, reaching 73.0% with value-model reranking.

Claim 41612 Label: fact Provenance: primary Recorded

Microsoft Research Blog

No stored spot-check names this claim in this edition.

Orchard-GUI, a 4-billion-parameter vision-language model, achieved an average success rate of 68.4% across the WebVoyager, Online-Mind2Web, and DeepShop benchmarks.

Claim 41614 Label: fact Provenance: primary Recorded

Microsoft Research Blog

No stored spot-check names this claim in this edition.

Assertion 13

- Moonshot AI's next useful signal is an official smaller deployment profile with measured power and operating cost. ChinaAI

Assertion status: No spot-check verdict is published for this assertion.

Deploying Moonshot AI's K3 model requires at least 16 H200 GPUs, a hardware cost that makes personal deployment unaffordable for most users.

Claim 41516 Label: fact Provenance: primary Recorded

ChinaAI

No stored spot-check names this claim in this edition.

Moonshot AI recommends deploying K3 on a super-node of at least 64 accelerator cards, with total costs of approximately 17 million RMB.

Claim 41517 Label: fact Provenance: primary Recorded

ChinaAI

No stored spot-check names this claim in this edition.

Running 64 K3 accelerator cards at full load consumes 45 kilowatts of electricity, exceeding the capacity of ordinary household electrical meters and wiring.

Claim 41518 Label: fact Provenance: primary Recorded

ChinaAI

No stored spot-check names this claim in this edition.

Assertion 14

- Circles should disclose cohort definitions and measurement periods behind its revenue and customer-loss gains. OpenAI

Assertion status: No spot-check verdict is published for this assertion.

Circles utilizes the OpenAI API and Codex to deliver AI-native telecommunications experiences.

Claim 41489 Label: fact Provenance: primary Recorded

OpenAI News

No stored spot-check names this claim in this edition.

The implementation of OpenAI technology by Circles resulted in a 22% increase in average revenue per user (ARPU).

Claim 41490 Label: fact Provenance: primary Recorded

OpenAI News

No stored spot-check names this claim in this edition.

The implementation of OpenAI technology by Circles resulted in a 9% reduction in customer churn.

Claim 41491 Label: fact Provenance: primary Recorded

OpenAI News

No stored spot-check names this claim in this edition.

Assertion 15

Princeton CITP Import AI Microsoft Research OpenAI ChinaAI

Assertion status: No spot-check verdict is published for this assertion.

Circles utilizes the OpenAI API and Codex to deliver AI-native telecommunications experiences.

Claim 41489 Label: fact Provenance: primary Recorded

OpenAI News

No stored spot-check names this claim in this edition.

The implementation of OpenAI technology by Circles resulted in a 22% increase in average revenue per user (ARPU).

Claim 41490 Label: fact Provenance: primary Recorded

OpenAI News

No stored spot-check names this claim in this edition.

The implementation of OpenAI technology by Circles resulted in a 9% reduction in customer churn.

Claim 41491 Label: fact Provenance: primary Recorded

OpenAI News

No stored spot-check names this claim in this edition.

Deploying Moonshot AI's K3 model requires at least 16 H200 GPUs, a hardware cost that makes personal deployment unaffordable for most users.

Claim 41516 Label: fact Provenance: primary Recorded

ChinaAI

No stored spot-check names this claim in this edition.

Moonshot AI recommends deploying K3 on a super-node of at least 64 accelerator cards, with total costs of approximately 17 million RMB.

Claim 41517 Label: fact Provenance: primary Recorded

ChinaAI

No stored spot-check names this claim in this edition.

Running 64 K3 accelerator cards at full load consumes 45 kilowatts of electricity, exceeding the capacity of ordinary household electrical meters and wiring.

Claim 41518 Label: fact Provenance: primary Recorded

ChinaAI

No stored spot-check names this claim in this edition.

Researchers from the University of Toronto, the Vector Institute, the University of Cambridge, and ServiceNow built a prototype computer virus that uses AI models to compromise computers and leverages their GPU resources for inference to infect more hosts.

Claim 41548 Label: fact Provenance: primary Recorded

Import AI

No stored spot-check names this claim in this edition.

The proof-of-concept worm operates using an open-weight LLM published in 2025 that fits on a single A100 GPU with 80GB of VRAM, without relying on monitored vendor APIs.

Claim 41550 Label: fact Provenance: primary Recorded

Import AI

No stored spot-check names this claim in this edition.

The prototype worm achieved approximately 80% success in vulnerability detection, 53% in exploitation, and 88% in self-replication, resulting in an overall attack success rate of about 37%.

Claim 41551 Label: fact Provenance: primary Recorded

Import AI

No stored spot-check names this claim in this edition.

Microsoft Research introduced Orchard, an open-source framework centered on Orchard Env, a Kubernetes-based environment service designed to support scalable agentic AI research across multiple task domains.

Claim 41611 Label: fact Provenance: primary Recorded

Microsoft Research Blog

No stored spot-check names this claim in this edition.

Orchard-SWE achieved a 69.7% score on the SWE-bench Verified benchmark using approximately 3 billion active parameters, reaching 73.0% with value-model reranking.

Claim 41612 Label: fact Provenance: primary Recorded

Microsoft Research Blog

No stored spot-check names this claim in this edition.

Orchard-GUI, a 4-billion-parameter vision-language model, achieved an average success rate of 68.4% across the WebVoyager, Online-Mind2Web, and DeepShop benchmarks.

Claim 41614 Label: fact Provenance: primary Recorded

Microsoft Research Blog

No stored spot-check names this claim in this edition.

A 2022 research report identified a critical privacy vulnerability in certain ballot scanning machines where the algorithm used to assign random numbers to electronic ballot records was deterministic and reversible, allowing the sequence of ballot casting to be reconstructed.

Claim 41646 Label: fact Provenance: primary Recorded

Princeton CITP - Freedom to Tinker

No stored spot-check names this claim in this edition.

Author Max Springer used AI coding agents and public records to reconstruct the casting order of 1.52 million in-person ballots (98.9% of the total) from Georgia's May 2026 primary election.

Claim 41647 Label: fact Provenance: primary Recorded

Princeton CITP - Freedom to Tinker

No stored spot-check names this claim in this edition.

In 114 of the 139 Georgia counties examined, the in-person casting order of ballots could be recovered from public files alone, enabling the identification of voters' ballots in specific contexts.

Claim 41648 Label: fact Provenance: primary Recorded

Princeton CITP - Freedom to Tinker

No stored spot-check names this claim in this edition.