To the line
Capabilities on the main line · PLANNING + TOOLS

Autonomous task execution

In one viewIn 2025–2026, models combined reasoning with search, code, files and application control. They can now receive a multi-step task, not just a question.

This begins the shift from assistant to agent. Today these systems still need human oversight: they accumulate errors, lose the goal on long tasks and learn little from their own experience.

StatusCURRENT
TypeCapabilities on the main line
Marker2026
Events in dossier15
Development chronology

Researched

2024

Reasoning as a separate mode

About this evento1 showed that extra inference-time compute can materially improve hard problem solving.

2024-11

Model Context Protocol

About this eventAnthropic opened a standard for connecting models to data and work tools.

2025-03

An agent platform

About this eventThe Responses API and Agents SDK combined reasoning, tools, orchestration and observability.

2025-03

Autonomous task horizon

About this eventMETR proposed measuring agents by the human task duration they complete at a given reliability.

2025-04

Agent2Agent

About this eventGoogle opened a protocol for agents from different vendors to exchange tasks and results.

2025-10

Designing agent workflows

About this eventAgentKit turned agent applications into an engineering layer with versions, evaluations and traces.

2026-06-09

Claude Fable 5

About this eventAnthropic released a model for long-running coding and knowledge-work tasks designed around hours or days of agentic work rather than a single answer.

2026-06-30

Claude Sonnet 5

About this eventAnthropic updated its mainstream model for coding, agent workflows and everyday knowledge work.

2026-07

Longer work in a product

About this eventChatGPT Work carries out longer tasks across files, apps and finished deliverables.

2026-09-01

Fable 5.1: sustained projects

About this eventA Claude update for multistep work; evaluations must account for fallback.

NEWverifiedPermanent page
2026-09-02

Gemini 3.8 Flash: agentic work

About this eventA GA release expands fast-model capabilities for coding and workflows.

NEWverifiedPermanent page
2026-09-03

GPT-6 Astra: computer use

About this eventOpenAI announced a new generation for complex work and application control. Access is phased.

NEWverifiedPermanent page

In progress

now

Long-horizon reliability

About this eventThe longer the action chain, the greater the chance of accumulating an error or misunderstanding the goal.

Planned

planned

Memory and learning from outcomes

About this eventAn agent needs to retain experience across tasks and improve without full model retraining.

Distant horizons

ahead

Verifiable autonomy

About this eventA system completes a long task independently while its plan, actions and outcome remain auditable.

Sources and research

Primary material behind this dossier: papers, lab publications and official reports.

Directions
Capabilities on the main line