DISPATCH

Everything that mattered in AI, one page a week.

Most AI news does not survive the week. This is the part that did — the releases, the research, and the shifts that actually change how we build. Designed & built to keep you up to date with things in AI without needing to be unemployed. Just refresh Saturday morning and review the last week's dispatch.

90 DISPATCHESWRITTEN EVERY FRIDAYNEXT UPDATE IN

DISPATCH 42

WEEK OF OCT 11 – 17, 2025

Cheap sub-agents, self-designed chips, and SKILL.md: the week the agent stack got two new layers

OpenAI contracted 10 gigawatts of its own accelerators, Anthropic pushed May's frontier coding performance to a third of the price and shipped a file format for agent expertise, and Nvidia's flagship chip finally reached US volume production.

Two levers moved in the same week. Anthropic cut near-frontier coding to one-third of its own Sonnet rate and shipped a filesystem format for packaging agent expertise, while OpenAI stopped depending entirely on other companies' chip designs. Both changes land on the same developer budget line: cost per reasoning step, and how much of your agent's behavior is code you can version.

The policy line kept running too. Meta agreed to let parents switch off teens' one-on-one AI chats, following the FTC's inquiry into AI companions and an August Reuters investigation into its chatbots' romantic conversations with minors. Microsoft, meanwhile, ended free security support for Windows 10 the same week it pushed Copilot deeper into Windows 11 — one mass-migration event used as an AI-adoption event.

The two silicon stories rhymed. On Monday OpenAI signed with Broadcom to design 10 gigawatts of its own accelerators; on Friday Jensen Huang signed the first Blackwell wafer produced on US soil at TSMC's Phoenix fab. Designing your own inference silicon on Monday and celebrating your supplier's chip onshored on Friday is the same race told from both ends.

FRI · Oct 17, 2025regulationmetasafetyftcchatbots

Meta lets parents switch off teens' AI chats, after the FTC inquiry

On October 17 Meta said parents will be able to disable their teens' one-on-one conversations with AI characters entirely, block specific characters, and see topic-level summaries of what their teens discuss — though not full transcripts. Meta's own AI assistant stays available with age-appropriate defaults. The controls land on Instagram first, in the US, UK, Canada and Australia, starting early next year.

It follows the FTC's inquiry into AI companions and minors and an August 2025 Reuters investigation that found Meta's AI rules permitted romantic and sensual conversations with minors. Earlier in the week Meta restricted teen Instagram accounts to PG-13 content by default and said the same rating logic would govern AI chats; OpenAI had shipped ChatGPT parental controls in late September after a lawsuit over a teenager's death.

Meta lets parents switch off teens' AI chats, after the FTC inquiry
image.cnbcfm.com

WHY IT MATTERS

Chatbot safety is becoming a shipped feature with an age-verification and topic-classification pipeline behind it, not a policy page. If your product has an assistant that talks to users without a verified age, the enforcement precedent is now set by the largest platforms plus an active FTC inquiry — and the design pattern they all landed on, supervised accounts plus topic-level reporting plus character-level blocking, is one you will be asked to match in an enterprise questionnaire.

FRI · Oct 17, 2025hardwarenvidiatsmcsupply chainblackwell

First US-made Blackwell wafer rolls off TSMC's Phoenix fab

Nvidia and TSMC marked the first Blackwell wafer produced at TSMC's Arizona facility in Phoenix on October 17, with Jensen Huang calling it the first time in recent US history that the single most important chip is manufactured on American soil. Nvidia said the milestone represents Blackwell reaching volume production; TSMC Arizona is slated to run 2-, 3- and 4-nanometer nodes as well as A16.

The wafer still requires advanced CoWoS packaging, which remains in Taiwan, so US-made does not yet mean a finished US-built GPU. The timing rhymes with the week's other silicon story: OpenAI agreed to design its own accelerators on October 13, four days before Nvidia's flagship reached US volume production.

First US-made Blackwell wafer rolls off TSMC's Phoenix fab
NVIDIA

WHY IT MATTERS

Accelerator supply is the one input your inference bill depends on that you cannot engineer around. Onshoring part of the Blackwell pipeline adds a second geography to a supply chain that had been single-sited, which matters for capacity planning and for any procurement story you tell a customer about where your compute comes from. The CoWoS packaging gap is the part to watch — it is still the bottleneck.

THU · Oct 16, 2025agentsanthropicagent infrastructureskillstooling

Anthropic ships Agent Skills: SKILL.md becomes the unit of agent expertise

Anthropic launched Agent Skills on October 16: a Skill is a directory containing a SKILL.md file with YAML frontmatter plus optional scripts, references and assets. The design is progressive disclosure — only the name and description of each installed Skill are preloaded into the system prompt, and the full instructions load into context only when the agent judges the Skill relevant to the task.

The same format works across Claude apps, Claude Code, the Claude API and the Claude Agent SDK. On the developer side, Skills can be attached to Messages API requests and a new /v1/skills endpoint handles programmatic versioning and management, with Skills requiring the Code Execution Tool beta to run bundled code. Launch partners included Box, Canva, Notion and Rakuten, alongside built-in skills for PowerPoint, Excel, Word and PDF files.

WHY IT MATTERS

This is the missing substrate between prompts and MCP. Prompts evaporate at session end and MCP servers are long-lived processes you have to run; a Skill is just files in a repository — reviewable in a pull request, diffed, shipped through the same CI as everything else. If your team's agent behavior currently lives in scattered system prompts, Skills give you a place to put it where it can be tested and rolled back.

WED · Oct 15, 2025modelsanthropicpricingagentscoding

Claude Haiku 4.5 puts May's frontier coding performance at one-third the price

Anthropic released Claude Haiku 4.5 on October 15 at $1 per million input tokens and $5 per million output — one-third the $3/$15 of Sonnet 4 and Sonnet 4.5 — while running more than twice as fast. It ships a 200,000-token context window, 64,000 max output tokens and a February 2025 knowledge cutoff, and scores 73.3% on SWE-bench Verified, roughly level with Claude Sonnet 4, GPT-5 and Gemini 2.5 from just months earlier.

Anthropic's stated deployment pattern is Sonnet 4.5 breaking a problem into multi-step plans and orchestrating a team of Haiku 4.5 sub-agents to execute subtasks in parallel. Availability was same-day across Claude apps, the Claude API, Amazon Bedrock, Google Vertex AI, Microsoft Foundry, and a GitHub Copilot public preview.

WHY IT MATTERS

This is a sub-agent pricing event, not a model-launch event. Once a reasoning step costs a third of what it did, the binding constraint on a 50-step agent loop stops being the model's capability and starts being tokens per call — so fan-out architectures that were wasteful last quarter become correct now. Budget your orchestration layer around per-call cost, because the cheap tier just became good enough to be the default worker.

MON · Oct 13, 2025computecustom siliconopenaibroadcominfrastructure

OpenAI signs Broadcom for 10 GW of OpenAI-designed accelerators

OpenAI and Broadcom announced a multi-year collaboration on October 13 to deploy 10 gigawatts of custom AI accelerators. OpenAI designs the chips and the systems; Broadcom develops and deploys them with an all-Ethernet rack fabric spanning switching, PCIe and optical interconnect. The companies signed a term sheet rather than a final supply contract, with rack deployments targeted to begin in the second half of 2026 and complete by the end of 2029.

It is the third leg of OpenAI's compute push, after the 6-gigawatt AMD Instinct agreement on October 6 and Nvidia's commitment of up to $100 billion for at least 10 gigawatts. Reuters put 10 GW at roughly the electricity demand of more than 8 million US households; Broadcom shares rose about 9-10% on the news.

WHY IT MATTERS

If you build on a frontier API, this is a bet about your unit economics, not OpenAI's balance sheet. Custom inference silicon co-designed with the model that runs on it is the most direct route to cutting cost per token, and it widens the gap between closed-API inference and whatever you could self-host. Architect for model substitution: keep the provider boundary thin so a cheaper token price does not force a rewrite.

DISPATCH 41

WEEK OF OCT 4 – 10, 2025

OpenAI turned ChatGPT into a runtime, AMD bet on it in its own shares, and Ant open-sourced a trillion parameters

OpenAI spent DevDay turning ChatGPT into a software platform, AMD agreed to sell it six gigawatts of silicon in exchange for equity in itself, the MPA told OpenAI that policing Sora 2 is OpenAI's job, and Ant Group put a trillion-parameter MoE on Hugging Face under MIT.

OpenAI spent Monday recasting ChatGPT as a software platform: an Apps SDK built as an extension of the Model Context Protocol, an AgentKit stack covering agent building, chat UI, connectors and evals, and Codex out of beta. Hours later AMD agreed to supply six gigawatts of Instinct GPUs while taking a large part of its upside in a warrant for its own shares.

The same Monday, the Motion Picture Association told OpenAI that preventing Sora 2 infringement is OpenAI's responsibility, not rightsholders' — the day OpenAI put Sora 2 into the developer API. Meanwhile the open-weights line kept compressing: Ant Group's inclusionAI published a trillion-parameter mixture-of-experts model under MIT, and Google gave developers a computer-use model that treats the browser as the API.

The through-line is where value is moving: to the host runtime that runs apps and agents, to second-source compute bought with equity, and to permissively licensed weights that reset the cost floor. Model launches, funding rounds and per-seat pricing are downstream of those three.

THU · Oct 9, 2025open weightsmoeant groupinference cost

Ant Group open-sources Ling-1T: 1T parameters, about 50B active, MIT license

Ant Group's inclusionAI published and open-sourced Ling-1T on 9 October: a trillion-parameter mixture-of-experts model that activates roughly 50 billion parameters per token at a 1/32 routing ratio, released under the MIT license. Context is 32K, extensible to 128K with YaRN, and it was pre-trained on more than 20 trillion tokens. The team says it is the largest FP8-trained foundation model to date, with FP8 mixed precision giving a 15%-plus end-to-end training speedup over BF16 at under 0.1% loss deviation, and reports 70.42% on AIME 2025 at around 4,000 output tokens per problem plus roughly 70% tool-call accuracy on BFCL V3 with only light instruction tuning.

It follows Ring-1T-preview, the trillion-parameter thinking model Ant open-sourced in September, and sits at the top of a three-line family — Ling non-thinking, Ring thinking, Ming multimodal. Weights are on Hugging Face and ModelScope, where the repository lists about 999.71 billion safetensors parameters and roughly 2 TB of weights; the maker's serving guidance assumes a multi-node cluster rather than a single box.

Ant Group open-sources Ling-1T: 1T parameters, about 50B active, MIT license
Hugging Face

WHY IT MATTERS

Sparse activation is the entire cost argument: Ling-1T carries trillion-parameter capacity at something like 50B-dense latency, and MIT means no per-token license fee — your marginal cost is cluster access, not vendor inference pricing. For data-residency or air-gapped workloads that is now a frontier-adjacent model you can run inside your own VPC, paid for in multi-node serving and about 2 TB of weights to move and store.

TUE · Oct 7, 2025google deepmindagentscomputer usegemini

Google ships a computer-use model, and the browser becomes the API

Google DeepMind released Gemini 2.5 Computer Use in preview on 7 October, available through the Gemini API in Google AI Studio and Vertex AI. It is a specialized version of Gemini 2.5 Pro post-trained for UI control, and the capability is exposed as a new computer_use tool meant to run in a loop: the user's request, a screenshot and a history of recent actions go in, and function calls — click, type, scroll, navigate a dropdown — come out. Google claims it beats leading alternatives across several web and mobile control benchmarks at lower latency, with evaluations run in-house and by Browserbase, whose headless browsers host Google's public demos.

The model card is unusually specific about scope: it is optimized for web browsers, is not yet tuned for OS-level control, shows only promise on mobile, and requires explicit end-user confirmation for certain actions such as making a purchase. Two days later, on 9 October, Google Cloud launched Gemini Enterprise, an agent platform priced from $21 to $30 per seat per month that wires in the Agent2Agent protocol, the Model Context Protocol and the Agent Payments Protocol.

Google ships a computer-use model, and the browser becomes the API
storage.googleapis.com

WHY IT MATTERS

Treating pixels as an interface contract makes the long tail of enterprise and legacy software with no API addressable, which changes the economics of RPA and makes UI regression testing automatable inside CI. Budget for what it adds: screenshot tokens and multi-step latency in the loop, a human confirmation gate on anything that moves money, and the fact that browser sessions, logins and authenticated state are now part of your agent's attack surface.

MON · Oct 6, 2025openaiagentsmcpdeveloper platforms

DevDay: ChatGPT becomes an app platform, and the Apps SDK rides on MCP

At DevDay in San Francisco on 6 October, OpenAI opened the Apps SDK in preview: third-party software can render interactive interfaces inside ChatGPT, with context shared through the Model Context Protocol. Pilot partners shipped the same day — Booking.com, Canva, Coursera, Expedia, Figma, Spotify and Zillow — and OpenAI framed the reach as over 800 million ChatGPT users. Apps went live to logged-in Free, Go, Plus and Pro users outside the EEA, Switzerland and the UK, with app submissions and a directory promised later this year.

The same keynote delivered AgentKit: a visual Agent Builder in beta, the embeddable ChatKit UI at general availability, an expanded Evals system with trace grading and prompt optimization, and a Connector Registry in beta for governing MCP servers across ChatGPT, Enterprise and the API. Codex reached general availability, GPT-5 Pro and Sora 2 entered the API, and gpt-realtime-mini shipped alongside. OpenAI says monetization will run through the Agentic Commerce Protocol, its open standard for checkout inside the chat.

WHY IT MATTERS

Orchestration, guardrails, evals and versioning are precisely the layers teams hand-roll on top of a model API, and OpenAI just bundled them at standard API model pricing — fast to ship on, harder to leave. Because the Apps SDK is MCP-native, your MCP server becomes the integration surface for the ChatGPT host and for other clients, so the lock-in worth tracking is the host layer, not the API; OpenAI's promise to publish an open implementation of the host environment is the item that would defuse it.

MON · Oct 6, 2025amdopenaicomputehardware

AMD pays in its own stock: 6 GW of Instinct for OpenAI, warrants for up to 160M shares

AMD and OpenAI announced a six-gigawatt agreement spanning multiple generations of Instinct GPUs, with the first one-gigawatt deployment of MI450 parts starting in the second half of 2026. AMD issued OpenAI a warrant for up to 160 million shares — roughly a tenth of the company — vesting on deployment milestones, on AMD share-price targets, and on OpenAI hitting the technical and commercial milestones needed to enable deployments at scale. Reported terms put the strike at one cent per share with a $600 ceiling, which is why the announcement added about $80 billion to AMD's market value.

Neither company would put a dollar figure on the deal, calling it worth billions. The timing is the sharper detail: it landed hours after OpenAI's DevDay keynote, extending a run in which OpenAI has committed to roughly a trillion dollars of compute build-out in a matter of weeks.

WHY IT MATTERS

For anyone buying inference, this is the first credible second source of frontier-scale accelerator capacity from late 2026, which makes MI450 a planning assumption rather than a hedge — and makes ROCm and vLLM portability work worth doing on your own schedule instead of under pressure. It also reprices compute commitments as balance-sheet instruments, since a supplier is taking equity for volume; expect that structure to shape multi-year availability, allocation and discounting across the market.

MON · Oct 6, 2025copyrightopenaisoraregulation

MPA to OpenAI: stopping Sora 2 infringement is your job, not the rightsholders'

On Monday the Motion Picture Association urged OpenAI to take immediate and decisive action against Sora 2 outputs that it says infringe its members' films, shows and characters. MPA CEO Charles Rivkin argued that OpenAI must acknowledge it remains their responsibility — not rightsholders' — to prevent infringement on the Sora 2 service, and that well-established copyright law applies. The statement followed Sam Altman's Friday blog post reversing Sora's copyright posture: rather than making studios opt out, OpenAI said it will give rightsholders more granular control over their characters and share video revenue with those who allow generation.

The collision was scheduled by OpenAI itself. Sora 2 launched on 30 September, was the top free app on the App Store within days, and on Monday morning — hours before the MPA's statement — OpenAI used DevDay to open Sora 2 to developers via the API. The failure mode was already public: clips featuring Mario, James Bond and other protected characters, which is what turned a product launch into an industry-liability question.

WHY IT MATTERS

This is the liability template for any product that generates from user-supplied IP. Opt-out pushes the enforcement cost onto rightsholders and reliably draws organized industry pushback; opt-in plus revenue sharing converts the same content into a licensed input with a metered price. If you ship generation features, the decision-relevant work is provenance tracking and per-character policy controls — we'll iterate is not a defense once a rights-holder trade association is the counterparty.

ARCHIVE

Go back in time

Every dispatch, newest first. Each week is written once and left as it was published.