DISPATCH

Everything that mattered in AI, one page a week.

Most AI news does not survive the week. This is the part that did — the releases, the research, and the shifts that actually change how we build. Designed & built to keep you up to date with things in AI without needing to be unemployed. Just refresh Saturday morning and review the last week's dispatch.

90 DISPATCHESWRITTEN EVERY FRIDAYNEXT UPDATE IN

DISPATCH 58

WEEK OF JAN 31 – FEB 6, 2026

Twenty minutes apart: Opus 4.6 and GPT-5.3-Codex

Anthropic shipped a million-token Opus and agent teams, OpenAI answered with a coding model that helped debug its own training, and the agent-harness discipline got its name.

Two frontier coding models landed twenty minutes apart, which is now the normal shape of a release week. What is worth reading is not who won the score table but what both shipped: better context management and cheaper long agent runs.

By Friday a practitioner essay had named the thing everyone was already doing — engineer the harness, not the prompt — and within days a frontier lab published its own version of the same argument.

THU · Feb 5, 2026Frontier modelsLong context

Claude Opus 4.6 brings a million-token context and agent teams

Anthropic upgraded its top model, targeting what it calls context rot — quality decay over long conversations. Opus 4.6 is the first Opus-class model with a 1M-token context window, supports up to 128k output tokens, and adds context compaction on the API to summarize and replace older context mid-task, plus adaptive thinking with effort controls.

In Claude Code it introduces agent teams as a research preview: multiple agents working in parallel on independent, read-heavy work such as codebase review. Anthropic reports the highest score on Terminal-Bench 2.0 and roughly +144 Elo over GPT-5.2 on GDPval-AA.

Claude Opus 4.6 brings a million-token context and agent teams
cdn.sanity.io

WHY IT MATTERS

For developers this is the long-context-and-orchestration release. Compaction plus 1M context directly addresses the context-management pain of long agent runs, and agent teams move harness design from one loop to a small supervised fleet.

THU · Feb 5, 2026Frontier modelsCoding agents

GPT-5.3-Codex, the first model OpenAI says helped build itself

Roughly twenty minutes after Anthropic's post, OpenAI released GPT-5.3-Codex, merging GPT-5.2-Codex's coding with GPT-5.2's reasoning. It runs 25% faster and uses less than half the tokens of its predecessor.

OpenAI says early versions debugged their own training runs, managed deployment and diagnosed evaluations — its first model instrumental in creating itself. Reported scores: Terminal-Bench 2.0 at 77.3% versus 64.0%, and OSWorld-Verified at 64.7% versus 38.2%. It is also the first OpenAI model rated High capability for cybersecurity under its Preparedness Framework, with API access held back pending safety review.

WHY IT MATTERS

Two things matter beyond the score table: token efficiency, which makes long agent runs cheaper, and mid-task steerability without losing context. The High cyber rating plus a delayed API rollout is the template for how frontier coding models will ship from here.

THU · Feb 5, 2026Agent platformsEnterprise

OpenAI launches Frontier, a control plane for agent fleets

Also on February 5, OpenAI announced Frontier: a platform for building, deploying and managing agents across enterprise systems, framed around giving agents the same things people need at work — shared business context, onboarding, feedback-driven evaluation, and scoped identity and permissions.

It is positioned as vendor-neutral, supporting third-party agents, and runs locally, in a customer cloud or hosted, with SOC 2 Type II and ISO 27001/27017/27018/27701 compliance. Early adopters include HP, Intuit, Oracle, State Farm, Thermo Fisher and Uber.

OpenAI launches Frontier, a control plane for agent fleets
TechCrunch

WHY IT MATTERS

This is agents-as-coworkers productized: identity, permissions, evals and audit trails wrapped around agent execution. If you are building internal agents, it is a direct competitor to a self-built harness plus orchestration layer.

THU · Feb 5, 2026HarnessesPractice

Mitchell Hashimoto names the discipline: engineer the harness

HashiCorp co-founder Mitchell Hashimoto published “My AI Adoption Journey,” a six-step account of adopting coding agents. Step five is the one that spread: every time an agent makes a mistake, engineer a permanent fix into its environment — a documentation file or a programmatic tool — rather than re-prompting, so the harness gets stronger as the agent works.

He pairs it with practical habits and a caveat that skill formation in junior engineers deeply worries him. The framing — agent equals model plus harness — was picked up within weeks by Thoughtworks' Birgitta Böckeler and echoed in OpenAI's own harness-engineering write-up.

WHY IT MATTERS

It gave practitioners vocabulary for what they were already doing, and the core claim is testable: with the same weights, harness changes move results more than model swaps in many workloads. The reactive rule — one observed failure, one permanent environmental fix — is cheap to adopt in any repo.

WED · Feb 4, 2026Open weightsVoice

Mistral ships a self-hostable realtime ASR model

Mistral released two speech-to-text models. Voxtral Realtime is 4B parameters under Apache 2.0 with weights on Hugging Face, using a purpose-built streaming architecture rather than chunking an offline model, with configurable latency down to sub-200ms. At 480ms delay it stays within 1–2% WER of the batch model.

Voxtral Mini Transcribe V2 is the closed batch model at $0.003/min with speaker diarization, word-level timestamps, context biasing and 13 languages, claiming roughly 3x the speed of ElevenLabs Scribe v2 at about a fifth of the price.

WHY IT MATTERS

A 4B open-weight streaming ASR model you can self-host on a single 16GB GPU changes what voice agents cost to run — realtime transcription stops being a paid API dependency and the audio stays inside your infrastructure.

MON · Feb 2, 2026FundingDeployment

Waymo raises $16B at a $126B valuation

Alphabet's Waymo closed a $16 billion round at $126 billion post-money — more than double its October 2024 valuation — led by Dragoneer, DST Global and Sequoia, with Alphabet remaining majority investor.

Waymo says it does more than 400,000 rides a week across six US metros, tripled annual volume to 15 million rides in 2025, and points to 127 million fully autonomous miles with a 90% reduction in serious-injury crashes. The capital funds groundwork for over 20 additional cities in 2026, including Tokyo and London.

WHY IT MATTERS

Robotaxis remain the clearest at-scale deployment of AI in the physical world, and this is the largest bet yet that the unit economics hold. Regulatory and city-by-city expansion, not model quality, is the binding constraint.

DISPATCH 57

WEEK OF JAN 24 – 30, 2026

Open weights closed the gap while the regulators moved

Kimi K2.5 brought parallel sub-agent swarms to open weights, a 30-person US lab pretrained a 400B model for $20M, and the week ended with a $3B copyright escalation.

Two open-weight releases this week mattered for different reasons. One pushed the capability ceiling and added a genuinely new orchestration idea. The other pushed the cost floor, and made the point that pretraining a frontier-class model from scratch is no longer a big-lab-only activity.

Meanwhile the legal and regulatory pressure kept building from both directions — a US criminal conviction for AI trade-secret theft and an EU probe into an AI feature baked into a social platform.

WED · Jan 28, 2026LitigationCopyright

Music publishers sue Anthropic for more than $3B

Universal Music Publishing Group, Concord and ABKCO — a coalition of 223 music publishers — filed a new copyright suit against Anthropic in the Northern District of California, alleging infringement of more than 20,000 songs with potential statutory damages exceeding $3 billion, which the plaintiffs call likely one of the largest non-class-action copyright cases filed in the US.

The complaint alleges Anthropic torrented more than 700 of the publishers' works from shadow libraries and names CEO Dario Amodei and co-founder Benjamin Mann individually. It expands a 2023 suit covering about 500 works, and follows Anthropic's $1.5B settlement over pirated books.

WHY IT MATTERS

The largest copyright escalation against a frontier lab yet, and a test of whether training-data claims reach outputs and individual executives rather than only the company's underlying datasets.

TUE · Jan 27, 2026Open weightsAgents

Kimi K2.5 ships open weights with parallel agent swarms

Moonshot AI released Kimi K2.5, built from Kimi K2 by continued pretraining over roughly 15T mixed visual and text tokens. It is natively multimodal and targets coding, vision and deep-research workloads.

The new idea is Agent Swarm: for complex tasks the model self-directs up to 100 sub-agents across up to ~1,500 parallel tool calls, which Moonshot says cuts execution time up to 4.5x versus a single agent. Independent evaluators at Vals ranked it the new #1 open-weight model, first among open models on 13 of 17 benchmarks.

Kimi K2.5 ships open weights with parallel agent swarms
www.marktechpost.com

WHY IT MATTERS

The strongest open-weight release of the month, and the first to ship sub-agent orchestration as a first-class capability rather than something you bolt on with a framework. Parallel fan-out is exactly where agent harnesses are heading.

TUE · Jan 27, 2026Open weightsEfficiency

A 30-person startup pretrains a 400B model for about $20M

Arcee AI released Trinity Large, a 400B-total / 13B-active sparse MoE trained on 17T tokens in roughly 30 days on 2,048 NVIDIA B300 GPUs, under Apache 2.0 with GGUF quantizations and a full technical report.

The routing is extreme — 256 experts with 4 active per token, about a 1.56% routing fraction — plus sliding-window attention at a 3:1 local-to-global ratio. Arcee claims 2–3x faster inference than peers in its weight class, and shipped raw checkpoints including a 10T-token base with no instruction data.

A 30-person startup pretrains a 400B model for about $20M
cdn.sanity.io

WHY IT MATTERS

This is the concrete data point on how cheap and reproducible large-scale pretraining has become. The raw base checkpoint is unusual openness for researchers.

TUE · Jan 27, 2026CopyrightLitigationTraining data

Anthropic bought millions of books, cut the spines off, scanned them and pulped the paper

The Washington Post published the first detailed account of more than 4,000 pages of court filings unsealed in Bartz v. Anthropic, the class action brought by novelist Andrea Bartz and the nonfiction writers Charles Graeber and Kirk Wallace Johnson. Among them was an internal planning document dated 13 April 2024: “Project Panama is our effort to destructively scan all the books in the world.” A second line in the same document reads: “We don't want it to be known that we are working on this.”

The mechanics were industrial. Anthropic began buying in early 2024 and hired Tom Turvey, who had spent years building the partnerships behind Google Books, to source stock — bulk lots from second-hand dealers, often in batches of tens of thousands, including Better World Books, World of Books and Wonder Book. A vendor proposal in the filings sought capacity to convert between 500,000 and two million books over six months. A hydraulic-powered cutter took the spine off each volume, high-speed production scanners captured the pages, and recycling companies collected the remaining paper. The filings put the spend at tens of millions of dollars within about a year.

Destroying the books is what made it lawful. Judge William Alsup held in June 2025 that scanning a book Anthropic had bought was fair use specifically because the print original was destroyed: “one replaced the other.” Keeping the paper alongside the scan would have meant two copies. The same order found that Anthropic's earlier download of more than seven million pirated books — from Books3, LibGen and the Pirate Library Mirror — was not fair use regardless of what the books were later used for, and that is the part the $1.5 billion settlement covered.

WHY IT MATTERS

The ruling rewards scan-and-destroy over scan-and-preserve and produces no public benefit whatsoever: the resulting corpus is private, and no reader, library or author can search it. Anthropic paid nothing for the legal half of this and $1.5 billion for the illegal half, which is a strange price list — the program that removed millions of physical books from the world was the free one. This story broke in January and drew a second, larger wave of coverage in August when further documents were unsealed; the ruling itself is from June 2025, so the arc runs across three separate news moments.

MON · Jan 26, 2026HardwareInference

Microsoft ships Maia 200, aimed at Nvidia's software moat

Microsoft announced Maia 200, its second-generation inference accelerator: TSMC 3nm, over 100 billion transistors, 216GB of HBM3e at 7 TB/s, 272MB of on-chip SRAM, native FP8/FP4, and an on-die NIC over standard Ethernet supporting clusters up to 6,144 accelerators.

Microsoft claims 30% better performance per dollar than the latest hardware in its fleet, 3x the FP4 performance of AWS's third-gen Trainium, and FP8 above Google's seventh-gen TPU. Deployment began in Azure US Central with a Maia SDK preview open to developers.

WHY IT MATTERS

A direct shot at Nvidia on the software side as well as silicon — custom inference hardware paired with an open SDK and PyTorch/Triton support, with Copilot and GPT-5.2 workloads moving onto it.

MON · Jan 26, 2026ScienceOpen weights

Nvidia open-sources a full weather and climate stack

At the American Meteorological Society meeting Nvidia launched Earth-2: open models, libraries and frameworks across the whole weather pipeline. Earth-2 Medium Range handles 15-day forecasts across 70+ variables on a new Atlas architecture; Nowcasting produces kilometre-resolution 0–6 hour storm forecasts; Global Data Assimilation builds initial atmospheric states in seconds on GPUs instead of hours on supercomputers.

Weights and code are on Hugging Face and GitHub via Earth2Studio. Nvidia says trained models run roughly 1,000x faster than conventional methods, and Israel's meteorological service reported a ~90% cut in compute time.

WHY IT MATTERS

A fully open end-to-end scientific pipeline, from data assimilation to local nowcasting, that moves weather prediction off supercomputer budgets — with real operational results from national weather agencies.

MON · Jan 26, 2026RegulationPolicy

The EU opens a DSA investigation into X over Grok

The European Commission opened a new formal Digital Services Act investigation into X and extended its existing proceedings on X's recommender systems. The new probe assesses whether X properly evaluated and mitigated systemic risks from deploying Grok's functionality in the EU, and whether a mandated ad-hoc risk assessment happened before deployment.

The Commission said the risks appear to have materialized, and extended the probe to cover X's switch to a Grok-based recommender. Potential breaches cited include Articles 34, 35 and 42.

WHY IT MATTERS

The first major EU test of whether shipping an AI feature inside a Very Large Online Platform triggers pre-launch risk-assessment duties under the DSA — a template that will apply well beyond X.

ARCHIVE

Go back in time

Every dispatch, newest first. Each week is written once and left as it was published.