Announced for complex work, coding and computer use, with access expanding in phases.
Chronicle
Every dated event on the line, newest first. Filter by branch or milestone.
From GPT and BERT to modern text, multimodal, image, video and audio models. The archive covers public releases and major research previews; every link points to the team's primary announcement, documentation or paper.
GA Flash release for coding, agents and complex workflows, with text output and multimodal input.
Meta updated its long-horizon agentic and coding model, available in Muse Code and Meta Model API.
An update for coding and long projects in paid Claude products and the API. Account for safeguards and fallback in evaluations.
Z.ai released a post-training update to the GLM-5.2 base with major gains in complex coding, terminal tasks and long-horizon agentic work.
Google released the GA version of a fast reasoning model with a one-million-token context and improvements for coding, web development and agentic workflows.
The final V4-Pro checkpoint superseded the preview, substantially improved agentic tasks and shipped with open weights under the MIT license.
Grok 4.6 improved long-running agents, coding and interactive product work, launching in Cursor, Grok Build, the API and partner platforms.
Meta released a local 30B agentic model under Apache 2.0 with image perception, tool use, failure recovery and official GGUF builds.
Meta updated its proprietary coding and agentic model Muse Spark, with version 1.2 improving multi-file development, computer use and long-running workflows.
Moonshot AI published Kimi K3 weights with instructions for Transformers, vLLM and SGLang. It is an infrastructure-scale open-weights release: 2.8T parameters does not imply an ordinary local run.
Black Forest Labs opened Early Access to a multimodal foundation for images, video, audio and action prediction, with individual capabilities rolling out in stages.
OpenAI's flagship strengthened agentic work in coding, biology and cybersecurity.
A million-context model focused on lower compute cost and faster prefill and decoding.
A compact reasoning model became the fast fallback for GPT-5.4 Thinking in ChatGPT.
A specialized agentic model combined Codex and GPT-5 training for long-running coding work.
The update improved subject consistency, vertical video and output up to 4K.
OpenAI unified fast answers, deeper reasoning and routing between modes in one product system.
A 20B MMDiT model focused on complex typography and precise text editing inside images.
Opus 4 and Sonnet 4 improved long-running coding, tool use and agentic workflows.
The video model added native sound, speech and ambience generation alongside visuals.
A new image generation improved detail, typography and speed.
Google expanded access to its music model and integrated it into creator tools.
An open family combined normal and thinking modes across dense and MoE sizes.
Reasoning models began reasoning over images while using the full ChatGPT tool set.
The natively multimodal Scout and Maverick MoE family expanded context and image understanding.
Google's thinking model combined reasoning, code and multimodal understanding with long context.
OpenAI's largest chat model became a research preview of scaled pretraining without a separate reasoning mode.
Anthropic combined fast responses and visible extended reasoning in one model.
A vision-language family learned documents, interfaces, charts and long video sequences.
An open reasoning model and its distillations showed reinforcement learning competing with closed reasoning systems.
An open MoE model with 671B total and 37B active parameters sharply reduced frontier-model cost.
Google emphasized native tool use, streaming multimodal interaction and future agents.
Sora moved from research preview into a product for generating and editing video.
The family added 11B/90B vision models and small 1B/3B text models for edge devices.
An open 0.5B-to-72B family improved multilingual, coding, math and structured-output capabilities.
The model spent dedicated compute before answering, improving math, code and complex reasoning.
A new image family paired the open Schnell model with more capable commercial variants.
Meta released a 405B model with 128K context and allowed its outputs to improve other models.
A mid-tier model surpassed the previous Opus on many tasks and improved coding and vision.
Google introduced a 1080p video model with cinematic styles and longer scenes.
A new Imagen generation improved photorealism, detail and visual consistency.
A single model handled text, vision and speech with low latency.
Open 8B and 70B models raised the quality of dialogue, reasoning and coding.
Haiku, Sonnet and Opus added vision and separated the family by speed, price and peak capability.
A new MMDiT architecture improved text rendering and adherence to complex prompts.
The model reached a million-token context across long documents, code, audio and video.
OpenAI framed long video generation as a step toward modelling visual-world dynamics.
An open mixture-of-experts model activated only part of its parameters per token, reducing inference cost.
Google introduced a family trained from the start across text, images, audio and video.
An open research-preview model extended image diffusion into short video.
A compact open model showed that architecture and data quality could compete with larger systems.
Image generation moved into ChatGPT with much stronger adherence to complex prompts.
Open weights became available for research and most commercial uses.
Claude gained a 100K context window, stronger coding and broad public availability.
The first public Claude offered long-form dialogue and a distinct approach to safety training.
OpenAI's flagship added image input, stronger reasoning and more reliable instruction following.
Meta released a 7B-to-65B research family and accelerated the open-weight movement.
A conversational interface and human-feedback training turned large language models into a mass-market product.
An open speech-recognition model combined multilingual transcription, translation and noise robustness.
An open model brought strong image generation to consumer GPUs and created a large fine-tuning ecosystem.
A large text encoder and cascaded diffusion models improved prompt fidelity.
One visual-language model learned many tasks from a few interleaved image-text examples.
Diffusion improved image resolution, realism and controllable editing.
A model trained on code translated natural language into programs and powered the first GitHub Copilot.
A large GPT-family model turned text instructions into generated images.
A shared image-text space enabled recognition of new visual categories without task-specific training.
At 175B parameters, in-context learning became a visible capability of a large language model.
Google reframed many language tasks as a single text-to-text problem.
Scaling a language model produced coherent long-form text and early convincing zero-shot transfer.
Bidirectional pretraining exposed both left and right context and reshaped search, classification and question answering.
The first GPT showed that one pretrained Transformer could be adapted to several language tasks.