DISPATCH

Everything that mattered in AI, one page a week.

Most AI news does not survive the week. This is the part that did — the releases, the research, and the shifts that actually change how we build. Designed & built to keep you up to date with things in AI without needing to be unemployed. Just refresh Saturday morning and review the last week's dispatch.

90 DISPATCHESWRITTEN EVERY FRIDAYNEXT UPDATE IN

DISPATCH 50

WEEK OF DEC 6 – 12, 2025

Open weights closed on the coding frontier, MCP went vendor-neutral, and OpenAI shipped GPT-5.2 at a higher price

Mistral's Devstral 2 put an open 123B model four points off the closed coding frontier and Anthropic handed MCP to the Linux Foundation, the same week OpenAI shipped GPT-5.2 at 40% higher input pricing, took $1B from Disney, and the White House moved to preempt state AI law.

Three of the week's five stories broke within hours of each other on Thursday December 11: OpenAI shipped GPT-5.2 to end its self-declared code red, Disney put $1 billion of equity into OpenAI while licensing 200-plus characters to Sora, and the White House signed Executive Order 14365 against state AI regulation. OpenAI's own launch page still carries no head-to-head Gemini 3 numbers.

Tuesday belonged to openness on both axes. Mistral's Devstral 2 put a 123B open model at 72.2% on SWE-bench Verified with a 24B sibling that runs locally, and Anthropic donated the Model Context Protocol to the Linux Foundation's new Agentic AI Foundation — the same protocol whose prompt-injection and authentication gaps were the year's recurring agent-tooling problem.

The week's other in-window items sit outside the model and tooling lanes: Nvidia's H200 export re-opening to China at a 25% US cut is hardware and trade rather than model news, and the OWASP Top 10 for Agentic Applications is a taxonomy release rather than an event.

THU · Dec 11, 2025frontier modelsopenaipricingevals

GPT-5.2 arrives to end OpenAI's code red — with a 40% input-price increase

OpenAI released GPT-5.2 on December 11 in Instant, Thinking and Pro tiers, roughly a month after GPT-5.1 and after CEO Sam Altman issued an internal code red memo redirecting resources to ChatGPT in response to Google's Gemini 3. The model carries a 400,000-token context window and an August 31, 2025 knowledge cutoff. OpenAI says GPT-5.2 Thinking beats or ties human professionals on 70.9% of comparisons on GDPval, its benchmark of knowledge-work tasks across 44 occupations, and that it sets a new state of the art on OpenAI's MRCRv2 long-context evaluation and at 98.7% on Tau2-bench Telecom for tool calling.

OpenAI's launch page lists no head-to-head numbers against Gemini 3; comparisons are against OpenAI's own predecessors, with competitive benchmarks against Gemini 3 Pro and Claude Opus 4.5 shown only during the press briefing. API input pricing rose to $1.75 per million tokens, about 40% above GPT-5.1, and Anthropic's Opus 4.5 still scores higher on SWE-bench Verified. Altman told CNBC that Gemini 3's impact on OpenAI's metrics was smaller than feared and that he expects to exit code red by January.

GPT-5.2 arrives to end OpenAI's code red — with a 40% input-price increase
cdn.arstechnica.net

WHY IT MATTERS

The capability number went up and so did the price, so re-run your cost model rather than your latency model: push bulk, batch and long-context work to cheaper or open-weight models and reserve GPT-5.2 Thinking for hard reasoning. The missing Gemini 3 comparison is the one your model-selection decision actually needs — vendor-run benchmarks citing OpenAI's own prior models are not a substitute for your own eval set on your own tasks.

THU · Dec 11, 2025licensinggenerative videocopyrightopenai

Disney buys $1B of OpenAI and licenses 200-plus characters to Sora

Disney agreed a three-year licensing deal giving Sora and ChatGPT Images access to more than 200 Disney, Marvel, Pixar and Star Wars characters — costumes, props, vehicles and environments included, but expressly excluding actor likenesses and voices. Disney is making a $1 billion equity investment in OpenAI and receives warrants to buy more; it also becomes a major OpenAI customer, deploying ChatGPT internally and building on OpenAI's APIs for Disney+, with a selection of fan-made Sora shorts to appear on the streaming service.

The deal landed the same day as GPT-5.2 and follows Disney's cease-and-desist to Google over alleged unlicensed use of its IP, with publisher copyright suits against OpenAI still running. Disney CEO Bob Iger framed it as opportunity rather than threat, saying he would rather participate in the growth than watch it happen and be disrupted by it. Sora and ChatGPT Images are expected to start generating the licensed characters in early 2026.

WHY IT MATTERS

Rights clearing for generative media is becoming character-level, negotiated and partially exclusive — so if your product renders licensed IP, assume rights-holders will gate the valuable assets behind named AI partners and demand guardrails plus provenance. That changes which models can legally generate what. Build your media pipeline behind a provider abstraction now, because your image or video vendor's license catalogue is now a product dependency that can expire.

THU · Dec 11, 2025regulationpolicypreemptioncompliance

Executive Order 14365 opens a federal campaign to preempt state AI law

President Trump signed Executive Order 14365, Ensuring a National Policy Framework for Artificial Intelligence, on December 11, declaring US policy to be a minimally burdensome national framework and directing the Attorney General to establish an AI Litigation Task Force to challenge state AI laws. It orders Commerce to publish an evaluation of conflicting state laws and to withhold non-deployment BEAD broadband funding from states that have them, and sets a March 11, 2026 deadline for an FTC policy statement treating state-mandated bias mitigation as a per se deceptive practice and for FCC examination of a federal AI reporting standard.

The order follows December 7 release of FY2026 NDAA text that dropped federal AI preemption after bipartisan pushback. Its final text is softer than the leaked draft — it no longer names California's SB 53 and expressly preserves state authority over child safety, data-center siting and state procurement. The order itself names Colorado's algorithmic-discrimination law as its example of a statute that may even force AI models to produce false results. Legal commentators note an executive order cannot by itself override state law; preemption requires Congress.

Executive Order 14365 opens a federal campaign to preempt state AI law
www.whitehouse.gov

WHY IT MATTERS

Nothing is preempted yet: state transparency and anti-discrimination regimes — Colorado's AI Act, California's SB 53 — remain on the books and enforceable while this litigates, so keep data handling and AI disclosure conservative rather than betting on federal cover. The dates that could actually change your compliance surface are the Commerce evaluation and FTC policy statement due March 11, 2026, plus whatever emerges from the DOJ task force — plan for two parallel regimes, not one.

TUE · Dec 9, 2025open weightscoding agentslicensingmistral

Mistral's Devstral 2 drags open weights to within four points of the closed coding frontier

Mistral released Devstral 2 on December 9: a 123B-parameter dense transformer with a 256,000-token context window that scores 72.2% on SWE-bench Verified, the execution-graded benchmark built from 500 real GitHub issues. It ships under a modified MIT license. Its 24B sibling, Devstral Small 2, scores 68.0% under Apache 2.0 and Mistral says it runs locally on consumer hardware; the company frames both as competitive with DeepSeek V3.2, which it says is five times larger than the 123B.

The models launched alongside Mistral Vibe, an Apache 2.0 terminal coding agent with file editing, ripgrep search, shell execution and MCP integration, with day-zero hooks into Cline, Kilo Code and Zed. Both checkpoints are free on Mistral's API through the launch period; posted pricing afterwards is $2.00 per million tokens for Devstral 2 and $0.30 for Small 2. Independent human evaluation through the Cline scaffold put Devstral 2's win rate over DeepSeek V3.2 at 42.8% against 28.6% losses.

WHY IT MATTERS

A 24B coding agent under Apache 2.0 that runs offline is an escape hatch from closed coding APIs: no per-token bill, no source leaving the machine, no vendor able to deprecate your model. The 123B's modified MIT is not MIT — it carries revenue terms worth reading before you ship it commercially. And treat 72.2% as a signal, not a guarantee: SWE-bench Verified is Python-heavy and drawn from public repos that may overlap training data.

TUE · Dec 9, 2025agent infrastructuremcpgovernanceprotocols

Anthropic hands MCP to the Linux Foundation as the Agentic AI Foundation launches

Anthropic donated the Model Context Protocol to the newly formed Agentic AI Foundation, announced December 9 as a directed fund under the Linux Foundation, co-founded by Anthropic, Block and OpenAI and supported by Google, Microsoft, AWS, Cloudflare and Bloomberg. MCP joins Block's goose and OpenAI's AGENTS.md as founding projects. The governing board sets strategy, budget and project admissions, while maintainers keep full autonomy over the specification and day-to-day decisions — the same stewardship model the Linux Foundation applies to Kubernetes, PyTorch and Node.js.

At donation, MCP reported more than 97 million monthly SDK downloads, more than 10,000 published servers, and first-class client support across ChatGPT, Claude, Cursor, Gemini, Microsoft Copilot and VS Code. Platinum members each contributed $350,000. The foundation stood up three working groups at launch: Protocol Evolution, Security (prompt injection and authentication standards) and Ecosystem (server certification and compatibility testing).

WHY IT MATTERS

The protocol your agent tooling speaks is now neutral infrastructure rather than one vendor's product: it changes through maintainer consensus, so no single company can unilaterally break compatibility or deprecate the transport your integrations depend on. The Security working group is the line item to watch — MCP's adoption ran ahead of its security maturity this year, and shared standards are how that gap closes rather than recurring per implementation.

DISPATCH 49

WEEK OF NOV 29 – DEC 5, 2025

Two permissive frontier models in 48 hours, while OpenAI called code red

DeepSeek and Mistral shipped 671B- and 675B-parameter open weights under MIT and Apache 2.0 on 1-2 December, the same week Altman froze OpenAI's ads and agent products to defend ChatGPT and Google pushed Gemini 3 Deep Think to paying subscribers.

AWS re:Invent ran in Las Vegas from 1 to 5 December, where Matt Garman's keynote introduced the Nova 2 model family, three long-running frontier agents, Trainium3 UltraServers, and new AgentCore controls for policy, memory and evaluation. In the same 72 hours the open-weight side of the industry shipped two near-frontier families and OpenAI declared an internal emergency.

The open-weights wave was the week's real story: DeepSeek's V3.2 on 1 December — 671B total, 37B active, MIT license, 128K context — and Mistral 3 on 2 December, with Mistral Large 3 at 675B total and 41B active under Apache 2.0 alongside 3B, 8B and 14B edge models. Both launches are explicitly pitched at agent workloads and tool use rather than chatbot benchmarks.

The closed frontier spent the week on defense and consolidation: Sam Altman's companywide memo freezing advertising and agent products to shore up ChatGPT, Google extending its Gemini 3 line to Ultra subscribers with Deep Think, Anthropic buying Bun — the runtime its coding agent already runs on — and OpenAI agreeing to buy training-observability vendor Neptune.

THU · Dec 4, 2025frontier modelsreasoningbenchmarksgooglepricing

Google ships Gemini 3 Deep Think to Ultra: 41.0% on Humanity's Last Exam

On 4 December Google rolled Gemini 3 Deep Think out to Google AI Ultra subscribers in the Gemini app, the reasoning mode it previewed when Gemini 3 Pro launched on 18 November. Google reports 41.0% on Humanity's Last Exam without tools and 45.1% on ARC-AGI-2 with code execution, attributing the jump to advanced parallel reasoning that explores multiple hypotheses simultaneously. Answers typically take minutes, and the mode is gated behind the $250/month Ultra tier.

The timing is pointed: this landed two days after Altman's code-red memo, marking the second Google release in three weeks to set the benchmark bar OpenAI is now spending a surge to match. It is also routed to consumers first — a toggle in the Gemini app — rather than to a documented API tier, so it is a demonstration of frontier reasoning capability more than a buildable endpoint.

Google ships Gemini 3 Deep Think to Ultra: 41.0% on Humanity's Last Exam
storage.googleapis.com

WHY IT MATTERS

The benchmark race is less interesting than the gating: when the top reasoning tier ships subscription-only, the buildable surface stays a generation behind at token prices. Plan your evaluation harness so swapping in a stronger model is a config change, and check whether your hardest tasks genuinely need frontier reasoning before assuming the next benchmark leader is the correct default for your traffic.

TUE · Dec 2, 2025open weightsapache 2.0mixture of expertsedge inferencemistral

Mistral 3: a 675B-parameter MoE under Apache 2.0, plus edge models you can run offline

Mistral announced the Mistral 3 family on 2 December: Mistral Large 3, a sparse mixture-of-experts with 41B active and 675B total parameters, trained from scratch on 3,000 NVIDIA H200 GPUs, plus three dense edge models at 3B, 8B and 14B. Every model ships under Apache 2.0 — Large 3 and the Ministral 3 series in base, instruct and reasoning variants. Mistral says Large 3 is its first MoE since Mixtral and debuts at #2 among non-reasoning open-weight models on LMArena, #6 across all open models.

Large 3 is multimodal for image understanding and multilingual, which Mistral claims is best-in-class for non-English, non-Chinese conversation; a reasoning variant is promised but not yet shipped. The Ministral tier is the cost half of the story: 3B to 14B models designed for offline and on-device use, fully fine-tunable, same permissive license. Coming one day after DeepSeek's MIT drop, that is a second permissive frontier family inside 48 hours.

Mistral 3: a 675B-parameter MoE under Apache 2.0, plus edge models you can run offline
Mistral AI

WHY IT MATTERS

Permissive weights at both ends of the size range redraw the build-versus-buy line: a 3B to 14B Apache 2.0 model you can fine-tune and ship on-device covers most extraction, classification and routing work without a network call, while a self-hostable 675B MoE is a credible fallback when a closed API's price, latency or availability is the risk you actually carry.

TUE · Dec 2, 2025openaicompetitionroadmapchatgptgoogle

Altman declares code red on ChatGPT, freezing ads and agent products

Sam Altman sent a companywide memo on Monday 1 December, reported by The Information and viewed by the Wall Street Journal, telling staff that OpenAI was declaring a code red on ChatGPT and that we are at a critical time for ChatGPT. The surge concentrates on speed, reliability, personalization and answering a wider range of questions, with a daily call for those responsible and encouragement to move teams temporarily.

Work on advertising, health and shopping agents, and the personal assistant codenamed Pulse is pushed back as a result. The trigger is Gemini 3, released 18 November, which beat OpenAI's models on industry benchmarks; on users, Google reported 650 million monthly Gemini users in October against ChatGPT's more than 800 million weekly actives. The echo is exact: Google declared its own code red in 2022 when ChatGPT launched.

WHY IT MATTERS

Two practical reads. First, if your roadmap depends on OpenAI shipping consumer agent surfaces — shopping, health, Pulse — treat those dates as having slipped, because capacity just moved to core chatbot quality. Second, a frontier lab publicly reallocating effort to product polish rather than capability is a signal to keep a second provider wired in behind your model router and to keep your evals provider-agnostic.

TUE · Dec 2, 2025agent toolingdeveloper toolsacquisitionjavascriptlicensing

Anthropic buys Bun, the JavaScript runtime its coding agent already runs on

Anthropic announced on 2 December that it is acquiring Bun — the JavaScript and TypeScript runtime, package manager, bundler and test runner created by Jarred Sumner in 2021 — disclosing in the same post that Claude Code passed $1 billion in run-rate revenue six months after its May 2025 general availability. Terms were not disclosed. Anthropic confirmed Bun stays open source under the MIT license and that it will keep investing in it as a general-purpose toolchain.

Bun was already load-bearing for Anthropic: it powers Claude Code's native installer and has been the runtime under the CLI since mid-2025. An agentic coding session is a hot loop — resolve dependencies, execute, run tests, read output, iterate — so owning the runtime means owning cold-start latency and install behavior that, at a billion dollars of run rate, are unit economics rather than developer ergonomics.

WHY IT MATTERS

The layer underneath your coding agent now belongs to a model lab. Expect Claude Code to get faster and more tightly integrated; the thing to watch is whether Bun's Node-compatibility and release cadence keep serving the general JavaScript ecosystem once its roadmap sits on one product's critical path. If you distribute a JS-based agent CLI, the license is safe but the incentives are no longer neutral.

MON · Dec 1, 2025open weightsagentssparse attentionlicensingdeepseek

DeepSeek ships V3.2 open weights: 671B parameters, MIT license, agents first

DeepSeek published the official DeepSeek-V3.2 family on 1 December: 671B total parameters with 37B activated per token, a 128K context window, and weights on Hugging Face under the MIT license, alongside a technical report. The only architectural change from V3.1-Terminus is DeepSeek Sparse Attention, which turns attention cost from quadratic to near-linear in sequence length and roughly halves the price of long-context inference — the cut DeepSeek first attached to its September experimental build. A second checkpoint, V3.2-Speciale, spends far more post-training compute and reaches gold-medal-level results at IMO 2025 and IOI 2025.

The report frames V3.2 as reasoning-first and built for agents: it is the first model in the V3 line to fold thinking directly into tool use, and its reinforcement-learning data comes from a large-scale agentic task-synthesis pipeline rather than chat traces alone. DeepSeek's own positioning is a daily driver at GPT-5-level performance — worth reading literally, since GPT-5 shipped in August, meaning open weights are now chasing a three-month-old closed target rather than this month's one.

WHY IT MATTERS

For anyone running agents, this is a license-clean checkpoint with thinking-in-tool-use support and long-context economics that stop punishing you per token: the decision it forces is whether your tool-calling loop should leave a paid API and live in your own VPC, where per-invocation cost is electricity instead of metered tokens.

ARCHIVE

Go back in time

Every dispatch, newest first. Each week is written once and left as it was published.