DISPATCH

Everything that mattered in AI, one page a week.

Most AI news does not survive the week. This is the part that did — the releases, the research, and the shifts that actually change how we build. Designed & built to keep you up to date with things in AI without needing to be unemployed. Just refresh Saturday morning and review the last week's dispatch.

90 DISPATCHESWRITTEN EVERY FRIDAYNEXT UPDATE IN

DISPATCH 86

WEEK OF AUG 15 – 21, 2026

Eight gigawatts, a $7.5B model router, and a training run put on hold

The money and megawatts moved faster than the models: OpenAI paused its largest frontier RL run, Stripe bought the routing layer, and DeepSeek added vision to its cheap agent stack.

No flagship US frontier model shipped this week. Instead, the infrastructure around models became the story — power, routing, safety overhead and the cost of seeing.

It is the week to remember that frontier AI is now a systems business before it is a model-release business.

FRI · Aug 21, 2026ModelsVision

DeepSeek adds vision to its cheap agent model and ships a free Files API

DeepSeek launched V4-Flash-Vision-Exp with image input, 1M context, JSON output, tool calls, Responses API and an Anthropic-compatible endpoint. Images cost roughly $0.17 per 1,000 at the published tokenization, while peak/off-peak pricing remains in place.

The free Files API lets you upload once and reference an image by file_id across requests. Harness 0.1.1 added support immediately.

WHY IT MATTERS

Screenshot-reading, chart parsing and self-correcting agent loops stop being a separate budget line when vision rides on the same inexpensive model.

TUE · Aug 18, 2026SafetyRL

OpenAI pauses its largest frontier RL run over its own cyber evals

OpenAI said it temporarily slowed scaling, including a two-week pause in deployment training, while its largest planned frontier RL run remained on hold. Preliminary evals suggested an upcoming model could meet the Critical cyber threshold; monitoring was extended to all tool-using inference.

Activation classifiers at every sampled token escalate to automated investigators with a 30-minute alert target. OpenAI put monitoring overhead at roughly 20% of observed inference compute.

OpenAI pauses its largest frontier RL run over its own cyber evals
img.helpnetsecurity.com

WHY IT MATTERS

This is the first public price tag attached to restraint: safety monitoring is not free, and the lab stopped its own biggest run because the evals said stop.

TUE · Aug 18, 2026ResearchAgents

Claude designs protein binders against 14 of 15 targets

Claude ran 24–48-hour autonomous campaigns against 16 protein targets: researching targets, choosing epitopes, installing and running open design models, and making every design decision without human input. Two independent labs synthesized every design exactly as delivered.

354 of 1,320 designs bound — a 26.8% hit rate — across 14 of 15 interpretable targets. On RBX1, 28 of 90 designs bound at 3.9 nM versus 9 of 245 human competition entries at 45 nM.

WHY IT MATTERS

An autonomous agent using open tools beat an expert design campaign on a task that used to take specialists weeks. The capability is useful, impressive and explicitly dual-use.

MON · Aug 17, 2026InfrastructureEnergy

OpenAI signs an 8-gigawatt Ohio data center

OpenAI secured roughly 8 GW at the former Portsmouth Gaseous Diffusion Plant in Ohio, with SB Energy building and operating it under a 20-year lease. Nvidia’s filing disclosed credit support capped at $105B for the first 4.25 GW and a $1.5B equity investment in SB Energy.

The site is exclusively Nvidia compute; the first 800 MW is expected in 2028. The deal includes community funds and Codex credits for Ohio students.

OpenAI signs an 8-gigawatt Ohio data center
image.cnbcfm.com

WHY IT MATTERS

This is the new frontier-compute template: the lab takes capacity, the chip vendor takes residual risk, and the contract pays its own grid upgrades.

MON · Aug 17, 2026AgentsOpen source

Nous Research ships Bot Mode, turning agent profiles into a roster of named bots

Nous Research shipped Bot Mode for Hermes Desktop, replacing the single-agent session list with a roster of named bots. Each bot is a real Hermes profile with its own chat, memory, skills, pinned model and avatar, and they message each other through a persistent Agent Inbox and hand work off by mention. It ships as a desktop plugin rather than a core patch, MIT-licensed, and is now bundled on by default.

The announcement drew real traction — past 900,000 views on the company's own post within a day, and reposted well beyond its usual circle. Hermes Agent is already on this timeline twice this quarter: a harness comparison on 7 August where it posted the lowest cost per success at $0.46, and a $1.5B valuation round reported on 13 July.

Hermes Bot Mode for the Hermes desktop app
Nous Research

WHY IT MATTERS

Most agent products sell you one assistant; this sells a team you assemble once, with separate memory and a separate model per role — which is the shape multi-agent work actually takes in practice. The per-bot persona-and-memory split is a better organising idea than one ever-growing chat.

DISPATCH 85

WEEK OF AUG 8 – 14, 2026

Four frontier releases in four days — and one proof you can actually check

Google, xAI, Meta and DeepSeek shipped into the developer tier while Anthropic moved a 167-year-old math bound and OpenAI made cyber capability a gated product.

The cheap-workhorse model moved again this week. Then it moved again. The useful competition is no longer just intelligence: it is price, context, harness compatibility and whether you can run the thing locally.

Off the leaderboard, Anthropic’s Riemann result was the week’s most consequential research because the proof is public and checkable.

THU · Aug 13, 2026ModelsPricing

Google ships Gemini 3.7 Flash at half the price

Gemini 3.7 Flash arrived three weeks after 3.6 with large gains on coding evals: FrontierCode 43.6% versus 34.4%, DeepSWE 65.3% versus 49%, and WebDev Arena Elo 1588 versus 1538. Its introductory price was half the previous Flash tier.

The release hit 939 points on Hacker News, a good proxy for the developer attention this slot now commands.

Gemini 3.7 Flash announcement
Google

WHY IT MATTERS

The cheap workhorse moved on both axes at once: capability and cost. If you pay per token for an agent harness, this should be in your next baseline comparison.

WED · Aug 12, 2026ModelsAgents

Grok 4.6 ties GPT-5.6 Sol on the third-party index

xAI released Grok 4.6 with a 500K context window and a $2/$6 per-million-token API price. It improved on Terminal-Bench, APEX-Agents, CursorBench and DeepSWE after a longer supplemental training run using model-generated reasoning data and agentic RL.

It is live in Cursor, Grok Build, OpenRouter, Vercel and Cloudflare — a distribution advantage that matters as much as the score.

WHY IT MATTERS

The mid-price agent tier is crowded enough that switching cost, harness compatibility and cache behavior matter more than a couple of leaderboard points.

TUE · Aug 11, 2026AgentsProducts

xAI opens Grok Bot in beta — always-on agents for everyone else

xAI opened Grok Bot in beta on 11 August for SuperGrok subscribers on desktop and iOS, with Windows and Linux builds out and Android promised. The framing is deliberately not a chatbot: each Bot gets its own cloud computer, signs into the tools and sites you already use, and carries a job to the end, returning only when something needs approval. The launch post's own line is that there is a huge difference between 90% done and 100% done.

It arrives through an unusual corporate route — a SpaceX rollup that folded xAI and Cursor into one company — and rides distribution rather than novelty: Grok was already reported at roughly 117 million monthly active users in March 2026.

Grok Bot announcement
xAI

WHY IT MATTERS

Agent work has been a developer story all year; this is the consumer version, and what it sells is completion rather than conversation. If you build agents, this is the behaviour your users will now expect — persistent, tool-connected, and quiet until it needs you.

MON · Aug 10, 2026Open weightsLocal inference

Meta open-weights a 30B agent model for one consumer GPU

Meta released Muse Glimmer under Apache 2.0, optimized for always-on local agents, function calling, coding and judge workflows. The three-phase recipe used teacher-logit distillation, longer-context agent data, SFT, on-policy distillation and RL.

Weights landed with developer docs and integrations for llama.cpp, MLX and ExecuTorch. A 30B agent model that runs locally changes the privacy and latency calculus for personal tooling.

WHY IT MATTERS

Meta’s distillation recipe is directly reusable: a large teacher can become a laptop-sized agent without giving up the function-calling interface developers actually need.

MON · Aug 10, 2026ResearchVerification

An unreleased Claude raises the Riemann bound from 41.6% to 67.2%

Anthropic reported an unreleased Claude combining results from separate papers to raise the proven lower bound on zeta zeros on the critical line from 41.6% to 67.2%. Claude also produced a publicly checkable Lean 4 proof; it did not prove the Riemann hypothesis.

Secondary reporting described roughly 36 hours of autonomous work across 60 subagents and 650 failed ideas. The important artifact is not the number, it is the formalization.

Anthropic mathematics research
Anthropic

WHY IT MATTERS

Long-horizon search is starting to move mathematics, not just benchmark scores. A Lean artifact lets anyone verify the claim instead of trusting a leaderboard.

ARCHIVE

Go back in time

Every dispatch, newest first. Each week is written once and left as it was published.