Skip to content
FLAVIO COPES
flaviocopes.com

Industry notes

AI news

This page is the database. Every AI note I track lives here: model launches, coding agent changes, and the date they shipped. The homepage lists the last five titles. My own product updates stay on News.

I add items by hand. I link the vendor post or the repo when I have one. This started on September 19, 2026.

Want new items in your feed reader? Subscribe to the AI news RSS feed.

~~~

70 notes

October 2026

OpenAI brought GPT-6 to everyone in ChatGPT, with Intelligent UI

GPT-6 is now the model behind the Chat tab in ChatGPT for everyone, free users included. Paid plans get GPT-6 Sol, while Free and Go get GPT-6 Luna. They replace GPT-5.6 Sol and GPT-5.6 Luna.

The big new thing is Intelligent UI. Answers aren’t just text anymore: they can include charts, interactive diagrams, buttons, forms, and small tools like a calculator or a bill splitter you use right in the conversation. GPT-6 can also start answering while it’s still thinking.

It started rolling out to Plus, Pro, Business, and Enterprise on October 7, and to Free and Go on October 8. The models used in ChatGPT Work and Codex don’t change with this release.

OpenAI announcement →

Anthropic released Claude Haiku 5.5

Claude Haiku 5.5 completes the Claude 5.5 family. Anthropic calls it its cheapest, fastest, and most capable small model, built for high-volume work like summaries, classification, browser use, and subagents working under Opus 5.5 or Sonnet 5.5.

It costs around 75% less to run than Haiku 4.5. For prompts up to 100k tokens it is $0.10 per million input tokens and $0.50 per million output tokens. It is also the first Haiku with an adjustable effort setting.

It is available today on the Claude Platform as claude-haiku-5-5, and on AWS, Google Cloud, and Microsoft Azure. Anthropic also cut Sonnet 5.5 cache-read prices in half, and is adding a monthly API credit for Max and Team subscribers.

Anthropic announcement →

Mistral released Mistral Large 4 in preview

Mistral Large 4 is Mistral’s new flagship model, and its largest so far. It has 1 trillion parameters, 49 billion of them active for each token, and it understands both text and images.

Mistral says it is the strongest open-weight model built outside China, with strong results on coding, agent work, cybersecurity, finance, and law. In a blind coding evaluation Mistral ran, it placed second of five models, behind only Claude Opus 5.

The preview is available today through the Mistral API in Mistral Studio. The weights come out by the end of October, so anyone will be able to download and run it.

Mistral announcement →

Earendil released Pi 1.0

Pi is a small, open source coding agent that runs in the terminal. It works with models from almost every provider, with a ChatGPT plan, or with a model running on your own computer. Earendil says hundreds of thousands of people use it every week, and calls 1.0 a stable version that people and businesses can depend on.

Pi used to leave MCP out on purpose. Now it connects to MCP servers, the common way agents talk to other apps, and the model can write a short script that calls several tools in one go. Extensions can also send each request to a different model, and the terminal screen has a new look and opens full screen.

Earendil also released Pi Durable, an experimental package for building agents that run for a long time and pick up where they left off after a crash.

Earendil announcement →A deep dive into Pi →

Cloudflare released Clef, an open decision model

Clef is Cloudflare’s first decision model, in the same family as TypeSafe’s Jev. Instead of writing text, it answers typed questions with probabilities, so software can route a ticket, block a request, or hand a case to a person.

It comes in two sizes, Clef and a faster Clef-flash, both hosted on Workers AI. The weights are open source on Hugging Face, it accepts the same requests as Jev, and unlike Jev it can also look at images. It costs more than Jev per request.

Cloudflare is also starting a service to fine-tune Clef on your own data, working hands-on with design partners first.

Cloudflare announcement →A deep dive into Clef →

September 2026

Google released Gemini 4 Argon

Gemini 4 Argon is Google’s new frontier model. Google DeepMind says it is strong on long-horizon software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.

It is rolling out first to trusted cyber defenders through Google’s Fairwind Program. Wider access for paid API customers and Google AI Ultra is planned after more guardrail work. Output context goes up to 1M tokens.

Introductory API pricing is $2 per million input tokens and $10 per million output tokens, with a steep discount on cached input. After the intro period it rises to $4 and $20.

Google announcement →

OpenAI released GPT-6.1 Sol

GPT-6.1 Sol is an upgrade to GPT-6 Sol. OpenAI says it nearly matches GPT-6 Astra on agentic coding, computer use, and professional work, at one-fifth of Astra’s standard API input and output prices.

It is available today in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu. It is not in Chat yet. In the API the model id is gpt-6.1-sol, at $2 per million input tokens, $0.10 per million cached input tokens, and $10 per million output tokens.

A faster Ultrafast option, up to 8x token generation in Codex, is coming in the next few days. Astra still leads on the hardest scientific tasks.

OpenAI announcement →

OpenAI launched dots

Dots are always-on agents from OpenAI, powered by GPT-6 Astra. Each dot has its own cloud computer and browser, can connect to thousands of apps, and keeps working toward your goals in the background.

You message or call your dot in ChatGPT, Slack, or Teams. It can also reach you with progress or a decision that needs approval. Conversations with the dot do not count toward ChatGPT usage limits. Codex and ChatGPT Work tasks it starts still do.

The first dot is included for Pro and Business Premium in eligible markets. Enterprise, Edu, and Healthcare can try the beta when an admin turns it on. You create it in the ChatGPT desktop app or desktop browser.

OpenAI announcement →A deep dive into OpenAI dots →

OpenAI launched the Decisions API

OpenAI launched the Decisions API for real-time decisions. You define questions with finite pre-defined answers, provide context as text or images, and get answers for classifying content, routing requests, or choosing an agent’s next action.

It is available in limited preview today, with a broad release planned in the coming days. This puts OpenAI in the same space as TypeSafe’s Jev, with a model making bounded decisions inside software instead of generating free-form chat.

OpenAI announcement →A deep dive into Jev →A deep dive into Clef →

Anthropic released Claude Sonnet 5.5

Claude Sonnet 5.5 is the second model in the Claude 5.5 family. Anthropic says it runs over 30% faster than Sonnet 5 and costs up to 30% less for most work, while scoring much higher on coding and knowledge-work benchmarks.

It is strongest at well-scoped everyday tasks, bug fixes, and polished documents, slides, and spreadsheets. Opus 5.5 stays the pick for complex open-ended work.

Sonnet 5.5 is available in Claude, Claude Code, the API, and on Amazon Bedrock, Google Cloud, and Microsoft Azure. Haiku 5.5 is still coming in the next few weeks.

Anthropic announcement →

OpenAI released GPT-6 Sol and GPT-6 Luna

GPT-6 Sol and GPT-6 Luna bring much of what makes GPT-6 Astra good into faster and cheaper models. Sol is meant for complex coding and agent work. Luna is the smallest and cheapest, for simple tasks you run a lot.

Both are rolling out in ChatGPT and Codex for paid plans, and in the API at half the price of the GPT-5.6 models. Free users can try Luna in the desktop app.

OpenAI announcement →

Anthropic released Claude Opus 5.5

Claude Opus 5.5 is the first model in the Claude 5.5 family. Anthropic says it performs close to Fable 5.1 on most work, while costing less and answering faster than Opus 5.

It is available in Claude, in Claude Code, and on the major cloud platforms. Sonnet 5.5 and Haiku 5.5 are coming in the next few weeks.

Anthropic announcement →

SpaceXAI released Grok 4.7

Grok 4.7 is the new SpaceXAI model for coding and longer knowledge work. It was trained on tasks that take many hours, and it knows how to work inside Grok Bot.

It costs the same as Grok 4.6, and there is a faster variant at a higher price. You can use it in Cursor, in Grok Build, and through the Grok API.

SpaceXAI announcement →A deep dive into Grok Bot →

Claude Code now reads AGENTS.md

Claude Code now reads AGENTS.md when a project has no CLAUDE.md file. If the project already has a CLAUDE.md, Claude Code keeps using that one.

You can change this behavior in the settings, including loading both files at once.

This is the file I already keep in repos so Cursor and Codex start with the same rules. Claude Code catching up means I can stop maintaining a second CLAUDE.md just for that tool.

agents-md plugin on GitHub →What is an AGENTS.md file →

TypeSafe left stealth with Jev

TypeSafe AI came out of stealth with $40M in funding and Jev, its first public model. Jev is not a chatbot and not a coding model. You give it data and a list of questions, and it answers each one with how confident it is.

TypeSafe sees it as a small, fast, cheap piece inside ordinary software, not something you chat with. Access is through a waitlist for now.

TypeSafe AI →A deep dive into Jev →

Cursor launched Projects

Cursor launched Projects for work that lasts longer than one chat. Each Project has a coordinator that plans the work, delegates it to other agents, and keeps shared context across local and cloud sessions.

The coordinator does not write code itself. It can also watch Slack, schedules, and pull requests, then start new work when one of those events happens.

Cursor announcement →How to use Cursor Projects →

OpenAI launched the Agents API beta

OpenAI now offers the engine behind Codex as an API, so developers can build their own long-running agents on top of it. OpenAI keeps the agent session alive and recovers it when something goes wrong.

The Agents API is in public beta for all developers.

OpenAI announcement →

OpenAI announced GPT-Live-1

GPT-Live-1 is a voice model that can listen while it talks, so the conversation feels more natural. When it needs to think or use a tool, it hands that work to another agent and keeps talking.

Voice is billed by the minute, and the work done by the other agent is billed separately.

OpenAI announcement →

Grok Bot gained Cursor cloud-agent orchestration

Every Grok Bot can now create Cursor cloud agents, read their transcripts, inspect screenshots attached as proof, and send follow-up instructions. A Bot can also interrupt a run that is going in the wrong direction.

This gives Grok Bot a way to find work outside a repository, then hand the coding part to Cursor.

Grok Bot for Engineering →Grok Bot vs Cursor Projects →

Meta launched Muse

Muse is a personal AI agent from Meta. You message it in the Muse app or in WhatsApp, and it can send email, book travel, and fill out forms. It keeps working after you close the app, and it asks before it sends an email or pays.

Meta runs it in its own protected cloud environment, and payments use a one-time card so your real card number stays hidden.

It is rolling out in the US, free up to a usage limit, with two paid plans for heavier use. This is the consumer agent, separate from Muse Glimmer, the open model Meta released in August.

Meta announcement →Muse subscription plans →

WebMCP added consequentialHint

WebMCP tools can now mark an action as consequential. Set consequentialHint for work that is significant or hard to reverse, such as booking a flight, transferring money, or deleting data.

The annotation defaults to false and remains a hint. The browser or agent still needs proper authorization and a confirmation step before it acts.

Announcement →WebMCP specification change →A deep dive into WebMCP →

Herdr added multi-machine agent management

Herdr 0.9 brings local and remote machines into the same terminal interface. You can add an SSH-reachable machine, see its agents beside your local ones, and move between them without opening another dashboard.

Each server still owns its own sessions and terminals. Herdr combines the view, not the underlying state.

Herdr announcement →A deep dive into Herdr →

OpenAI released GPT-6 Astra

GPT-6 Astra is the new OpenAI flagship model. It started with a small group of companies, then rolled out to paid ChatGPT plans and the big cloud platforms.

OpenAI presents it as the step after GPT-5.6 Sol for coding, browsing, using a computer, and professional work. Because it is so capable at hacking tasks, the most advanced security features stay restricted.

I have not used Astra as my daily driver yet. Sol is still the model I send precise coding work to in Cursor.

OpenAI announcement →

Anthropic released Claude Fable 5.1

Fable 5.1 is an update to Fable 5, at the same price.

Anthropic says it is better on long agent runs and less likely to declare success when it is stuck. That matches how I use Fable: planning, review, and judgment, not the first pass at a small code change.

Anthropic announcement →

August 2026

Grok Bot opened to more Cursor and SuperGrok plans

Ten days after launch, xAI expanded Grok Bot beyond SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium. SuperGrok Plus, Cursor Pro+, and every Cursor Teams subscriber got in, plus a limited trial for other users.

You still sign in with a Cursor account. The Bot still runs on one cloud computer shared by all your Bots.

xAI announcement →A deep dive into Grok Bot →

Cursor launched Origin code hosting

Cursor entered code hosting with Origin. It includes repositories, pull requests, code browsing, agents, and a staged early-beta rollout across paid plans.

An Origin-hosted repository uses Origin as its source of truth. You can also sync an existing GitHub repository while GitHub remains the source of truth.

Cursor announcement →A deep dive into Cursor Origin →

SpaceX closed the Cursor acquisition

SpaceX closed the Cursor deal on August 14, 2026. Cursor sits in the same company as Grok Bot. That is why a Grok Bot login is a Cursor login, and why a Bot can hand work to a Cursor cloud agent.

xAI had already taken the SpaceXAI name in July. The August close is the one that put the coding IDE and the always-on Bot under one roof.

The Verge →Grok Bot vs Cursor Projects →

Z.ai released GLM-5.3

GLM-5.3 is Z.ai's coding model for long-running engineering work. It can keep a very large amount of code in view at once.

It was available first through the GLM Coding Plan. Two weeks later Z.ai released the model so anyone can download and run it.

Z.ai announcement →A deep dive into ZCode →

Zed announced Delta

Delta is a separate coding environment from Zed. DeltaDB keeps agent conversations and worktrees together, so you can follow the work while it happens and review the reasoning beside the code.

The August announcement opened a private beta. Delta entered public beta on September 16. It works with existing Git repositories, and normal commits and pushes still work.

Zed announcement →Public beta update →

xAI launched Grok Bot

Grok Bot entered public beta on August 11, 2026 for SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium. A Bot is a named worker with its own memory, skills, and routines. It runs on a persistent cloud computer that keeps going after you close the laptop.

All of your Bots share that one machine. They share files, browser sessions, and credentials. I use it for jobs that start behind a login, not for work that should end in a pull request.

xAI announcement →A deep dive into Grok Bot →

Meta released Muse Glimmer

Muse Glimmer is an open model from Meta for agents that run on your own computer. It can read text and images, use tools, and recover when a tool fails.

You can download it for free, including smaller versions that fit on a well-equipped laptop.

Meta announcement →A deep dive into open-weight AI models →

Cloudflare introduced Kitesurf

Kitesurf is a web browser built for AI agents, running on Cloudflare instead of on Chrome. Existing browser automation tools can talk to it without changes.

It is in beta. It uses far fewer resources than a normal browser, but it is slower and does not work with every website yet.

Cloudflare announcement →A deep dive into Kitesurf →

bb launched as a programmable IDE for coding agents

bb is an open-source, local-first IDE and orchestrator for coding agents. It launched with support for Codex, Claude Code, Cursor, and ACP agents while using the subscriptions you already have.

The interface is programmable too. You can ask an agent to create a plugin or add a feature to bb itself.

bb announcement →How bb is built →

Cloudflare introduced @cloudflare/computer

Cloudflare Computer is an open-source project that gives an AI agent its own files and a place to run code, and keeps them around between sessions.

It is an early preview, with a lightweight option and a full Linux option.

Cloudflare announcement →A deep dive into Cloudflare Computer →

July 2026

Moonshot launched Kimi K3

Kimi K3 is Moonshot's new large model. It understands text and images and can work with very long inputs.

Moonshot launched it as a hosted service first, then released the model for download later in July.

Moonshot announcement →

LM Studio launched Bionic

Bionic is a separate app from LM Studio for coding and document work. It can use a local model, a model on another machine through LM Link, or LM Studio Secure Cloud.

Code projects can inspect repositories, edit files, search code, and show inline diffs. Work projects add a sandbox, checkpoints, previews, and web search.

LM Studio announcement →

GPT-5.6 Sol became generally available

OpenAI previewed the GPT-5.6 family on June 26, 2026: Sol as the flagship, Terra for everyday work, Luna as the cheap fast one. A small partner group got API and Codex access first.

General availability landed on July 9, 2026 in ChatGPT, Codex, and the API. Sol is the model I still pick when the task is specified well enough that I want the code written carefully.

OpenAI announcement →

xAI became SpaceXAI

On July 6, 2026 the xAI account on X switched to @SpaceXAI and shipped a new logo. SpaceX had already bought xAI in February. July was the public name change, not the ownership change.

Grok, X, and later Cursor all sit under that brand. Product pages still say xAI in a lot of places, including the Grok Bot docs.

Engadget →

June 2026

Vercel launched eve

eve is Vercel's open-source framework for building and running AI agents. An agent is a directory with Markdown instructions and TypeScript tools, plus optional skills, subagents, channels, and schedules.

Durable sessions, sandboxed execution, approvals, tracing, and evals come with the framework. The first release was a public preview.

Vercel announcement →Vercel eve: an open framework for building AI agents →

Anthropic released Claude Fable 5

Fable 5 is a new tier above Opus. Anthropic called it a Mythos-class model made safe enough for general use. It costs about twice as much as Opus.

Mythos 5 is the same model with some safeguards off, available only through Project Glasswing. On the public product, some Fable 5 requests can still get handed to Opus 5.

This is the model that made the named-tier thing real for me. Opus stayed in the mix. Fable became the one I use when the job is judgment, not typing.

Anthropic announcement →

May 2026

April 2026

OpenAI released GPT-5.5

GPT-5.5 rolled out in ChatGPT and Codex, with a faster mode in Codex.

Developers got it through the API one day later.

OpenAI announcement →

OpenAI expanded Codex beyond coding

Codex gained background computer use, image generation, and more than 90 plugins. OpenAI also added GitHub review comments, multiple terminals, an in-app browser, and SSH development boxes in alpha.

Automations could resume an existing thread and keep working over days or weeks. Memory launched as a preview.

OpenAI announcement →The complete guide to Codex →

March 2026

Cursor released self-hosted cloud agents

Cursor cloud agents can now run on your own machines instead of Cursor servers. Your machine connects to Cursor on its own, so there is nothing to open up in your network.

Cursor still handles the models and the planning, while the commands run on your machine. There is also a setup for companies that want to run many of them.

Cursor announcement →Run Cursor cloud agents on your own Mac mini →

February 2026

Anthropic released Claude Opus 4.6

Claude Opus 4.6 was the first Opus model that could work with very large projects in one go, as a beta feature.

The same release introduced Agent Teams in Claude Code as a research preview. One lead agent could split a job among several teammates and combine their work.

Anthropic announcement →What I learned reading Claude's system prompts →

OpenAI released GPT-5.3-Codex

GPT-5.3-Codex was built for long-running coding and research tasks. OpenAI said it ran 25% faster than GPT-5.2-Codex and could keep working across larger projects.

It launched across the Codex app, CLI, IDE extension, and cloud for paid ChatGPT plans.

OpenAI announcement →The complete guide to Codex →

OpenAI launched the Codex app

The first Codex desktop app launched on macOS. It put several agents in one workspace, let them work in parallel, and included skills plus scheduled Automations.

This was the standalone Codex app. OpenAI later moved Codex into the ChatGPT desktop app.

OpenAI announcement →The complete guide to Codex →

SpaceX acquired xAI

SpaceX completed its acquisition of xAI on February 2, bringing xAI, X, Starlink, and SpaceX into the same company.

SpaceX described orbital data centers as a long-term goal. The company later adopted the SpaceXAI name and bought Cursor.

SpaceX announcement →

January 2026

Moonshot released Kimi K2.5

Kimi K2.5 is a Moonshot model for coding and tool use that understands both text and images.

It also introduced Agent Swarm, where one model coordinates several agents working on parts of the same job.

Moonshot announcement →

November 2025

Anthropic released Claude Opus 4.5

Claude Opus 4.5 launched in Claude, the API, and the major cloud platforms, at a much lower price than earlier Opus models.

Claude Code gained an editable planning workflow with the same release. You could change the plan before the agent started writing code.

Anthropic announcement →

Google launched Gemini 3 and Antigravity

Google released Gemini 3 Pro through the Gemini app, Search, AI Studio, and Vertex AI. The same announcement introduced Antigravity, a coding environment built around agents.

Antigravity entered a free public preview. Its agents could work across the editor, terminal, and browser.

Google announcement →Choosing between an IDE, CLI, and vibe coding tool →

OpenAI released GPT-5.1

GPT-5.1 launched in ChatGPT as Instant and Thinking. Both models adjusted how much reasoning they used to match the task.

The ChatGPT rollout started on November 12. API access followed one day later.

OpenAI announcement →

September 2025

August 2025

July 2025

Vercel released AI SDK 5

AI SDK 5 added typed chat support for React, Svelte, Vue, and Angular. It also expanded tool handling, agent-loop controls, and speech support.

This was the main TypeScript framework release from Vercel that summer. The current AI SDK has moved several major versions beyond it.

Vercel announcement →A deep dive into the Vercel AI SDK →

Z.ai released GLM-4.5

GLM-4.5 is a free, downloadable model from Z.ai for coding and agent work. It can answer quickly or take time to think first.

Z.ai offered it both as a hosted service and as a download. Later GLM releases replaced it in ZCode.

Z.ai announcement →A deep dive into ZCode →

OpenAI launched ChatGPT agent

ChatGPT agent combined deep research, Operator, a visual browser, a terminal, and connected apps. It could research a task, use websites, run code, and take actions from the same conversation.

This release also absorbed the standalone Operator workflow. Deep research remained available as a separate feature.

OpenAI announcement →How AI agents should handle passwords →

Moonshot released Kimi K2

Kimi K2 is a large model from Moonshot for coding and tool use, released for free download.

An updated version followed in September. Kimi K2 Thinking and Kimi K2.5 came later.

Moonshot announcement →

xAI released Grok 4

Grok 4 launched with built-in tool use and live search, in the Grok app and the API. xAI also introduced Grok 4 Heavy, a pricier version where several agents work on the same problem.

Grok 4 Fast followed in September, then Grok 4.1 in November.

xAI announcement →

May 2025

Anthropic made Claude Code generally available

Claude Code reached general availability three months after its research preview. Anthropic added IDE integrations, background tasks, and broader access for developers.

The release arrived with Claude Opus 4 and Sonnet 4, which became the new models behind Claude Code.

Anthropic announcement →Choosing between an IDE, CLI, and vibe coding tool →

April 2025

OpenAI released o3 and o4-mini

OpenAI released o3 and o4-mini as reasoning models that could use ChatGPT tools while they worked. They could browse the web, run Python, inspect uploaded files, and reason over images.

Both models later changed availability, so this entry describes their launch state.

OpenAI system card →

OpenAI launched Codex CLI

The first Codex CLI was an open-source terminal agent installed through npm. It could inspect a repository, edit files, and run commands from the terminal.

OpenAI later replaced most of that original implementation with a Rust codebase. The product kept the Codex CLI name.

First Codex CLI commit →The complete guide to Codex →

OpenAI released GPT-4.1

OpenAI launched GPT-4.1 in three sizes for developers. All of them could read very large inputs, like whole codebases.

The family focused on coding and following instructions carefully. It was only available through the API at launch.

OpenAI announcement →

February 2025

Anthropic released Claude Code in research preview

Anthropic introduced Claude Code as a terminal coding agent alongside Claude 3.7 Sonnet. It could inspect a codebase, edit files, run commands, and work with Git from the terminal.

The first release was a limited research preview. Claude Code reached general availability three months later, on May 22.

Anthropic announcement →Choosing between an IDE, CLI, and vibe coding tool →

xAI announced Grok 3 Beta

Grok 3 Beta included Grok 3, Grok 3 mini, and Think variants that spent more time reasoning. xAI also introduced DeepSearch for researching information across the web.

The models were still in beta, and API access was only promised at launch.

xAI announcement →

GitHub previewed Copilot agent mode

Copilot agent mode could edit several files, inspect its own errors, run another iteration, and suggest terminal commands. The first preview required VS Code Insiders.

This was the point where Copilot moved beyond inline suggestions and chat toward a coding-agent loop inside the editor.

GitHub announcement →

January 2025

OpenAI released o3-mini

o3-mini was a small, fast reasoning model. You could choose how long it thought before answering.

It focused on text and code, and it could not read images.

OpenAI announcement →

DeepSeek released DeepSeek-R1

DeepSeek released R1 as an open-weight reasoning model with website and API access. The weights and code used the MIT license.

The release also included six smaller distilled models. DeepSeek compared R1 with leading closed reasoning models using its own evaluations.

DeepSeek announcement →