Coding agent benchmarks
Usage shows what people choose; benchmarks show what an agent can finish. This page tracks Terminal-Bench, the public leaderboard whose entries are complete agents (the harness plus the model driving it) working through hard, long tasks in a real terminal, with the cost of each run where it is published.
Terminal-Bench 4.0
Leaderboard updated Sep 21, 2026 · 27 runs from 4 agents
| # | Agent | Best run: model | Tasks solved | Cost of run | Runs listed |
|---|---|---|---|---|---|
| 1 | CodexOpenAI |
GPT-6 Astra · max | 58.2% ± 2.8 | $3,267 | 8 |
| 2 | Claude CodeAnthropic |
Fable 5.1 · max | 57.9% ± 3.8 | $6,244 | 14 |
| 3 | Grok BuildxAI |
Grok 4.7 · xhigh | 37.6% ± 3.5 | $3,683 | 3 |
| 4 | mini-SWE-agentSWE-agent |
Gemini 3.8 Flash · high | 19.1% ± 3.4 | $1,829 | 2 |
All 27 runs on Terminal-Bench 4.0
| Agent | Model | Effort | Tasks solved | Cost of run | Model date |
|---|---|---|---|---|---|
| Codex | GPT-6 Astra | max | 58.2% ± 2.8 | $3,267 | Sep 3, 2026 |
| Claude Code | Fable 5.1 | max | 57.9% ± 3.8 | $6,244 | Sep 1, 2026 |
| Codex | GPT-6 Astra | xhigh | 57.9% ± 2.7 | $2,351 | Sep 3, 2026 |
| Codex | GPT-6 Astra | high | 57.9% ± 3.0 | $2,269 | Sep 3, 2026 |
| Claude Code | Fable 5.1 | xhigh | 57.9% ± 3.4 | $4,872 | Sep 1, 2026 |
| Claude Code | Fable 5.1 | high | 54.5% ± 3.4 | $3,985 | Sep 1, 2026 |
| Codex | GPT-6 Astra | medium | 54.2% ± 2.7 | $1,915 | Sep 3, 2026 |
| Claude Code | Fable 5.1 | medium | 53.9% ± 3.4 | $2,833 | Sep 1, 2026 |
| Claude Code | Opus 5 | xhigh | 53.9% ± 3.2 | $6,086 | Jul 24, 2026 |
| Claude Code | Opus 5 | max | 51.8% ± 3.4 | $5,969 | Jul 24, 2026 |
| Codex | GPT-6 Astra | low | 50.6% ± 2.8 | $1,557 | Sep 3, 2026 |
| Claude Code | Opus 5 | high | 50.3% ± 3.7 | $4,662 | Jul 24, 2026 |
| Claude Code | Opus 5 | medium | 44.9% ± 3.8 | $3,192 | Jul 24, 2026 |
| Claude Code | Fable 5 | max | 44.5% ± 3.9 | $7,265 | Jun 9, 2026 |
| Claude Code | Fable 5.1 | low | 43.3% ± 3.6 | $2,359 | Sep 1, 2026 |
| Claude Code | GLM-5.3 | max | 41.8% ± 3.2 | $2,728 | Aug 14, 2026 |
| Grok Build | Grok 4.7 | xhigh | 37.6% ± 3.5 | $3,683 | Sep 21, 2026 |
| Codex | GPT-5.6 Sol | max | 37.3% ± 3.8 | $2,542 | Jun 26, 2026 |
| Claude Code | Opus 5 | low | 34.9% ± 3.9 | $2,394 | Jul 24, 2026 |
| Claude Code | Opus 4.8 | max | 23.6% ± 3.6 | $6,481 | May 28, 2026 |
| Codex | GPT-5.6 Terra | max | 21.5% ± 3.2 | $1,734 | Jun 26, 2026 |
| Grok Build | Grok 4.6 | high | 20.3% ± 3.1 | $3,592 | Aug 12, 2026 |
| mini-SWE-agent | Gemini 3.8 Flash | high | 19.1% ± 3.4 | $1,829 | Sep 2, 2026 |
| Codex | GPT-5.6 Luna | max | 17.3% ± 2.9 | $347 | Jun 26, 2026 |
| Grok Build | Grok 4.5 | high | 12.4% ± 2.6 | $2,094 | Jul 16, 2026 |
| Claude Code | Sonnet 5 | max | 12.4% ± 3.1 | $9,604 | Jun 30, 2026 |
| mini-SWE-agent | Gemini 3.7 Flash | high | 11.2% ± 2.5 | $1,262 | Aug 13, 2026 |
Which agents improved most on Terminal-Bench 4.0
Each agent's best run with the earliest model it was listed with, against its best run to date. The gain comes from newer models and from changes to the agent itself; the board does not separate the two.
| Agent | First listed run | Tasks solved | Best run to date | Tasks solved | Gain |
|---|---|---|---|---|---|
| Claude Code | Opus 4.8 · May 28, 2026 | 23.6% | Fable 5.1 · Sep 1, 2026 | 57.9% | 34.2 pts |
| Grok Build | Grok 4.5 · Jul 16, 2026 | 12.4% | Grok 4.7 · Sep 21, 2026 | 37.6% | 25.2 pts |
| Codex | GPT-5.6 Sol · Jun 26, 2026 | 37.3% | GPT-6 Astra · Sep 3, 2026 | 58.2% | 20.9 pts |
Terminal-Bench 3.0
Leaderboard updated Sep 3, 2026 · 12 runs from 6 agents
| # | Agent | Best run: model | Tasks solved | Cost of run | Runs listed |
|---|---|---|---|---|---|
| 1 | mini-SWE-agentPrinceton |
Opus 5 · max | 42.7% ± 3.1 | $5,818 | 1 |
| 2 | CodexOpenAI |
GPT-5.6 Sol · max | 34.6% ± 3.1 | $3,951 | 3 |
| 3 | Claude CodeAnthropic |
Fable 5 · max | 34.0% ± 3.4 | $6,481 | 5 |
| 4 | Grok BuildxAI |
Grok 4.6 · high | 26.5% ± 2.9 | $2,095 | 1 |
| 5 | DevinCognition |
SWE-1.7 Lightning | 18.6% ± 3.0 | $7,215 | 1 |
| 6 | Cursor CLICursor |
Grok 4.5 · xhigh | 15.7% ± 2.9 | $766 | 1 |
All 12 runs on Terminal-Bench 3.0
| Agent | Model | Effort | Tasks solved | Cost of run | Model date |
|---|---|---|---|---|---|
| mini-SWE-agent | Opus 5 | max | 42.7% ± 3.1 | $5,818 | Jul 24, 2026 |
| Codex | GPT-5.6 Sol | max | 34.6% ± 3.1 | $3,951 | Jul 13, 2026 |
| Claude Code | Fable 5 | max | 34.0% ± 3.4 | $6,481 | Jul 13, 2026 |
| Claude Code | GLM 5.3 | max | 32.4% ± 2.9 | $1,771 | Aug 18, 2026 |
| Grok Build | Grok 4.6 | high | 26.5% ± 2.9 | $2,095 | Aug 12, 2026 |
| Claude Code | Opus 4.8 | max | 21.1% ± 3.0 | $5,214 | Jul 12, 2026 |
| Codex | GPT-5.6 Terra | max | 20.8% ± 2.7 | $2,485 | Jul 12, 2026 |
| Devin | SWE-1.7 Lightning | — | 18.6% ± 3.0 | $7,215 | Aug 18, 2026 |
| Cursor CLI | Grok 4.5 | xhigh | 15.7% ± 2.9 | $766 | Jul 15, 2026 |
| Claude Code | Sonnet 5 | max | 14.6% ± 2.9 | $6,891 | Jul 12, 2026 |
| Codex | GPT-5.6 Luna | max | 14.3% ± 2.5 | $1,598 | Jul 12, 2026 |
| Claude Code | GLM 5.2 | max | 4.6% ± 1.9 | $3,402 | Jul 19, 2026 |
Terminal-Bench 2.1
Leaderboard updated Sep 3, 2026 · 22 runs from 6 agents
| # | Agent | Best run: model | Tasks solved | Cost of run | Runs listed |
|---|---|---|---|---|---|
| 1 | CodexOpenAI |
GPT-6 Astra · high | 87.4% ± 1.8 | $774 | 8 |
| 2 | Claude CodeAnthropic |
Fable 5 · xhigh | 83.8% ± 2.3 | $553 | 5 |
| 3 | Terminus 2Terminal-Bench |
Fable 5 · high | 80.5% ± 2.3 | $439 | 5 |
| 4 | Cursor CLICursor |
Grok 4.5 · high | 79.3% ± 2.9 | $134 | 1 |
| 5 | mini-SWE-agentPrinceton |
Muse Spark 1.1 · xhigh | 76.2% ± 2.4 | $198 | 1 |
| 6 | Gemini CLIGoogle |
Gemini 3 Pro · high | 65.8% ± 2.7 | $248 | 2 |
All 22 runs on Terminal-Bench 2.1
| Agent | Model | Effort | Tasks solved | Cost of run | Model date |
|---|---|---|---|---|---|
| Codex | GPT-6 Astra | high | 87.4% ± 1.8 | $774 | Sep 3, 2026 |
| Codex | GPT-6 Astra | medium | 87.0% ± 1.9 | $619 | Sep 3, 2026 |
| Codex | GPT-6 Astra | low | 86.7% ± 1.8 | $463 | Sep 3, 2026 |
| Codex | GPT-6 Astra | max | 86.7% ± 1.4 | $1,070 | Sep 3, 2026 |
| Codex | GPT-6 Astra | xhigh | 85.8% ± 1.4 | $889 | Sep 3, 2026 |
| Claude Code | Fable 5 | xhigh | 83.8% ± 2.3 | $553 | Jun 9, 2026 |
| Codex | GPT-5.5 | xhigh | 83.2% ± 2.2 | $2,059 | Apr 23, 2026 |
| Terminus 2 | Fable 5 | high | 80.5% ± 2.3 | $439 | Jun 9, 2026 |
| Cursor CLI | Grok 4.5 | high | 79.3% ± 2.9 | $134 | Jul 8, 2026 |
| Claude Code | Opus 4.8 | high | 78.9% ± 2.6 | $287 | May 28, 2026 |
| Codex | GPT-5.6 Terra | max | 78.4% ± 2.5 | $421 | Jun 26, 2026 |
| Terminus 2 | GPT-5.5 | xhigh | 78.0% ± 2.4 | $494 | Apr 23, 2026 |
| mini-SWE-agent | Muse Spark 1.1 | xhigh | 76.2% ± 2.4 | $198 | Jul 9, 2026 |
| Codex | GPT-5.6 Luna | max | 75.7% ± 2.6 | $241 | Jun 26, 2026 |
| Claude Code | Sonnet 5 | high | 74.6% ± 3.2 | $288 | Jun 30, 2026 |
| Terminus 2 | Gemini 3 Pro | high | 73.9% ± 2.5 | $224 | Nov 18, 2025 |
| Claude Code | Opus 4.7 | max | 68.9% ± 2.8 | $600 | Apr 16, 2026 |
| Terminus 2 | Opus 4.7 | max | 66.1% ± 2.7 | $582 | Apr 16, 2026 |
| Gemini CLI | Gemini 3 Pro | high | 65.8% ± 2.7 | $248 | Nov 18, 2025 |
| Gemini CLI | Gemini 3.1 Pro | high | 65.8% ± 3.3 | $236 | Feb 19, 2026 |
| Terminus 2 | Gemini 3.1 Pro | high | 65.6% ± 3.2 | $230 | Feb 19, 2026 |
| Claude Code | GLM-5.1 | max | 58.6% ± 2.4 | $277 | Mar 27, 2026 |
Which agents improved most on Terminal-Bench 2.1
Each agent's best run with the earliest model it was listed with, against its best run to date. The gain comes from newer models and from changes to the agent itself; the board does not separate the two.
| Agent | First listed run | Tasks solved | Best run to date | Tasks solved | Gain |
|---|---|---|---|---|---|
| Claude Code | GLM-5.1 · Mar 27, 2026 | 58.6% | Fable 5 · Jun 9, 2026 | 83.8% | 25.2 pts |
| Terminus 2 | Gemini 3 Pro · Nov 18, 2025 | 73.9% | Fable 5 · Jun 9, 2026 | 80.5% | 6.5 pts |
| Codex | GPT-5.5 · Apr 23, 2026 | 83.2% | GPT-6 Astra · Sep 3, 2026 | 87.4% | 4.3 pts |
Terminal-Bench 2.0
Leaderboard updated Aug 28, 2026 · 142 runs from 43 agents
| # | Agent | Best run: model | Tasks solved | Runs listed |
|---|---|---|---|---|
| 1 | NexAU-AHEchina-qijizhifeng |
GPT-5.5 | 84.7% ± 2.1 | 1 |
| 2 | LemonHarnessLR AILab of Lenovo CTO Org |
Multiple · v1.0.0 | 84.5% ± 2.6 | 2 |
| 3 | CapyCapy |
GPT-5.5 | 83.2% ± 2.1 | 2 |
| 4 | Codex CLIOpenAI |
GPT-5.5 · v0.121.0 | 82.2% ± 2.2 | 7 |
| 5 | PolarisPolarisOps |
Multiple · v2.2.0 (Polaris) | 82.2% ± 2.8 | 1 |
| 6 | TongAgentsBIGAI |
Gemini 3.1 Pro · v0.8.0 | 80.2% ± 2.6 | 2 |
| 7 | WOZCODEWOZCODE |
Claude Opus 4.7 | 80.2% ± 2.1 | 1 |
| 8 | SageAgentOpenSage |
GPT-5.3-Codex | 78.4% ± 2.2 | 2 |
| 9 | DroidFactory |
GPT-5.3-Codex | 77.3% ± 2.2 | 5 |
| 10 | Meta-HarnessStanford IRIS |
Claude Opus 4.6 · v1.1.0 | 76.4% ± 2.4 | 1 |
| 11 | CodeBrain-1.5Feeling AI |
GPT-5.3-Codex · v1.0.0 | 75.8% ± 2.0 | 2 |
| 12 | Codeliakousw |
GPT-5.3-Codex | 75.7% ± 2.2 | 1 |
| 13 | Simple CodexOpenAI |
GPT-5.3-Codex | 75.1% ± 2.4 | 1 |
| 14 | Terminus-KIRAKRAFTON AI |
Gemini 3.1 Pro · v3.3.0 | 74.8% ± 2.6 | 2 |
| 15 | MuxCoder |
GPT-5.3-Codex | 74.6% ± 2.5 | 4 |
| 16 | MAYA-V2ADYA |
Claude 4.6 Opus · v1.5-official | 72.1% ± 2.2 | 2 |
| 17 | spoox-o-mTUM |
GPT-5.3-Codex | 71.5% ± 2.5 | 3 |
| 18 | Junie CLIJetBrains |
Multiple · v507.2 | 71.0% ± 2.9 | 2 |
| 19 | AnteAntigma Labs |
Gemini 3 Pro | 69.4% ± 2.1 | 1 |
| 20 | IndusAGI Coding AgentVarun Israni (SoloVpx) |
GPT-5.3-Codex | 69.1% ± 2.3 | 2 |
| 21 | CruxRoam |
Claude Opus 4.6 | 66.9% | 5 |
| 22 | Deep AgentsLangChain |
GPT-5.2-Codex · v0.0.1 | 66.5% ± 3.1 | 1 |
| 23 | clnkrclnkr |
GPT-5.5 · v0.3.11 | 66.1% ± 2.5 | 1 |
| 24 | Terminus 2Terminal-Bench |
GPT-5.3-Codex · v2.0.0 | 64.7% ± 2.7 | 33 |
| 25 | II-AgentIntelligent Internet |
Gemini 3 Pro | 61.8% ± 2.8 | 1 |
| 26 | hookeleDmitry Barakhov |
GPT-5.1-Codex-Mini · v0.0.1 | 61.6% ± 1.9 | 1 |
| 27 | Gemini CLIGoogle |
Gemini 3.1 Pro · v0.35.0 | 61.4% ± 4.1 | 5 |
| 28 | WarpWarp |
Multiple | 61.2% ± 3.0 | 3 |
| 29 | Letta CodeLetta |
Claude Opus 4.5 · vlatest | 59.1% ± 2.4 | 3 |
| 30 | Abacus AI DesktopAbacus.AI |
Multiple · v1.106.24701 | 58.4% ± 2.8 | 1 |
| 31 | Claude CodeAnthropic |
Claude Opus 4.6 · v2.1.34 | 58.0% ± 2.9 | 5 |
| 32 | Grok CLISuperagent |
Grok 4.20 Reasoning · v1.1.1 | 57.3% | 1 |
| 33 | GooseBlock |
Claude Opus 4.5 · vstable | 54.3% ± 2.6 | 3 |
| 34 | Simplai AgentSimplAI |
Claude Sonnet 4.6 · v0.3.0 | 53.4% ± 2.8 | 1 |
| 35 | OpenHandsOpenHands |
Claude Opus 4.5 · v1.1.0 | 51.9% ± 2.9 | 12 |
| 36 | OpenCodeAnomaly Innovations |
Claude Opus 4.5 | 51.7% | 1 |
| 37 | CAMEL-AICAMEL-AI |
Claude Sonnet 4.5 · v1.0 | 46.5% ± 2.4 | 1 |
| 38 | Harness AgentlazyFrogLOL |
MiniMax M2.7 Highspeed | 42.9% ± 2.9 | 1 |
| 39 | cchuterteamblobfish.com |
minimax-m2.5 | 42.7% ± 2.8 | 1 |
| 40 | Mini-SWE-AgentPrinceton |
Claude Sonnet 4.5 | 42.5% ± 2.8 | 13 |
| 41 | Dakou Agentiflow |
Qwen 3 Coder 480B · v1.18.72-pre | 27.2% ± 2.6 | 1 |
| 42 | little-coderItay Inbar |
Qwen3.6-35B-A3B · v0.1.14 | 24.6% ± 3.2 | 3 |
| 43 | Bash AgentUCSB-SURFI |
TermiGen-32B · v1.0.0 | 19.3% ± 2.0 | 1 |
All 142 runs on Terminal-Bench 2.0
| Agent | Agent version | Model | Effort | Tasks solved | Model date |
|---|---|---|---|---|---|
| NexAU-AHE | — | GPT-5.5 | — | 84.7% ± 2.1 | Apr 23, 2026 |
| LemonHarness | 1.0.0 | Multiple | — | 84.5% ± 2.6 | May 14, 2026 |
| Capy | — | GPT-5.5 | — | 83.2% ± 2.1 | Apr 23, 2026 |
| Codex CLI | 0.121.0 | GPT-5.5 | — | 82.2% ± 2.2 | Apr 23, 2026 |
| Polaris | 2.2.0 (Polaris) | Multiple | — | 82.2% ± 2.8 | May 14, 2026 |
| TongAgents | 0.8.0 | Gemini 3.1 Pro | — | 80.2% ± 2.6 | Feb 19, 2026 |
| WOZCODE | — | Claude Opus 4.7 | — | 80.2% ± 2.1 | Apr 16, 2026 |
| LemonHarness | 1.0.0 | Multiple | — | 79.9% ± 3.0 | May 14, 2026 |
| SageAgent | — | GPT-5.3-Codex | — | 78.4% ± 2.2 | Feb 5, 2026 |
| Droid | — | GPT-5.3-Codex | — | 77.3% ± 2.2 | Feb 5, 2026 |
| Meta-Harness | 1.1.0 | Claude Opus 4.6 | — | 76.4% ± 2.4 | Feb 5, 2026 |
| CodeBrain-1.5 | 1.0.0 | GPT-5.3-Codex | — | 75.8% ± 2.0 | Feb 5, 2026 |
| Codelia | — | GPT-5.3-Codex | — | 75.7% ± 2.2 | Feb 5, 2026 |
| Capy | — | Claude Opus 4.6 | — | 75.3% ± 2.4 | Feb 5, 2026 |
| Simple Codex | — | GPT-5.3-Codex | — | 75.1% ± 2.4 | Feb 5, 2026 |
| Terminus-KIRA | 3.3.0 | Gemini 3.1 Pro | — | 74.8% ± 2.6 | Feb 19, 2026 |
| Terminus-KIRA | 3.3.0 | Claude Opus 4.6 | — | 74.7% ± 2.6 | Feb 5, 2026 |
| Mux | — | GPT-5.3-Codex | — | 74.6% ± 2.5 | Feb 5, 2026 |
| MAYA-V2 | 1.5-official | Claude 4.6 Opus | — | 72.1% ± 2.2 | Feb 5, 2026 |
| TongAgents | 0.7.0 | Claude Opus 4.6 | — | 71.9% ± 2.7 | Feb 5, 2026 |
| spoox-o-m | — | GPT-5.3-Codex | — | 71.5% ± 2.5 | Feb 5, 2026 |
| Junie CLI | 507.2 | Multiple | — | 71.0% ± 2.9 | Mar 7, 2026 |
| Droid | — | Claude Opus 4.6 | — | 69.9% ± 2.5 | Feb 5, 2026 |
| Ante | — | Gemini 3 Pro | — | 69.4% ± 2.1 | Nov 18, 2025 |
| IndusAGI Coding Agent | — | GPT-5.3-Codex | — | 69.1% ± 2.3 | Feb 5, 2026 |
| Crux | — | Claude Opus 4.6 | — | 66.9% | Feb 5, 2026 |
| Deep Agents | 0.0.1 | GPT-5.2-Codex | — | 66.5% ± 3.1 | Jan 14, 2026 |
| Mux | — | Claude Opus 4.6 | — | 66.5% ± 2.5 | Feb 5, 2026 |
| clnkr | 0.3.11 | GPT-5.5 | — | 66.1% ± 2.5 | Apr 23, 2026 |
| SageAgent | — | Gemini 3 Pro | — | 65.2% ± 2.1 | Nov 18, 2025 |
| Droid | — | GPT-5.2 | — | 64.9% ± 2.8 | Dec 11, 2025 |
| Terminus 2 | 2.0.0 | GPT-5.3-Codex | — | 64.7% ± 2.7 | Feb 5, 2026 |
| Junie CLI | 507.2 | Gemini 3 Flash | — | 64.3% ± 2.8 | Dec 17, 2025 |
| Droid | — | Claude Opus 4.5 | — | 63.1% ± 2.7 | Nov 24, 2025 |
| Codex CLI | 0.73.0 | GPT-5.2 | — | 62.9% ± 3.0 | Dec 11, 2025 |
| Terminus 2 | 2.0.0 | Claude Opus 4.6 | — | 62.9% ± 2.7 | Feb 5, 2026 |
| CodeBrain-1.5 | 1.0.0 | Gemini 3 Pro | — | 62.2% ± 2.6 | Nov 18, 2025 |
| II-Agent | — | Gemini 3 Pro | — | 61.8% ± 2.8 | Nov 18, 2025 |
| hookele | 0.0.1 | GPT-5.1-Codex-Mini | — | 61.6% ± 1.9 | Nov 12, 2025 |
| Gemini CLI | 0.35.0 | Gemini 3.1 Pro | — | 61.4% ± 4.1 | Feb 19, 2026 |
| Warp | — | Multiple | — | 61.2% ± 3.0 | Dec 12, 2025 |
| Droid | — | Gemini 3 Pro | — | 61.1% ± 2.8 | Nov 18, 2025 |
| Mux | — | GPT-5.2 | — | 60.7% | Dec 11, 2025 |
| Codex CLI | 0.63.0 | GPT-5.1-Codex-Max | — | 60.5% ± 2.7 | Nov 19, 2025 |
| Gemini CLI | 0.34.0 | Gemini 3.1 Pro | — | 59.4% ± 4.2 | Feb 19, 2026 |
| Letta Code | latest | Claude Opus 4.5 | — | 59.1% ± 2.4 | Nov 24, 2025 |
| Warp | — | Multiple | — | 59.1% ± 2.8 | Nov 20, 2025 |
| Mux | — | Claude Opus 4.5 | — | 58.4% | Nov 24, 2025 |
| Abacus AI Desktop | 1.106.24701 | Multiple | — | 58.4% ± 2.8 | Dec 11, 2025 |
| Claude Code | 2.1.34 | Claude Opus 4.6 | — | 58.0% ± 2.9 | Feb 5, 2026 |
| Crux | — | GPT-5.1-Codex | — | 57.8% ± 2.9 | Nov 12, 2025 |
| Terminus 2 | 2.0.0 | Claude Opus 4.5 | — | 57.8% ± 2.5 | Nov 24, 2025 |
| Grok CLI | 1.1.1 | Grok 4.20 Reasoning | — | 57.3% | Mar 10, 2026 |
| Terminus 2 | 2.0.0 | Gemini 3 Pro | — | 56.9% ± 2.5 | Nov 18, 2025 |
| Letta Code | latest | Gemini 3 Pro | — | 56.0% ± 3.0 | Nov 18, 2025 |
| Goose | stable | Claude Opus 4.5 | — | 54.3% ± 2.6 | Nov 24, 2025 |
| Terminus 2 | 2.0.0 | GPT-5.2 | — | 54.0% ± 2.9 | Dec 11, 2025 |
| Letta Code | latest | GPT-5.1-Codex | — | 53.5% ± 2.8 | Nov 12, 2025 |
| Simplai Agent | 0.3.0 | Claude Sonnet 4.6 | — | 53.4% ± 2.8 | Feb 17, 2026 |
| Terminus 2 | 2.0.0 | GLM 5 | — | 52.4% ± 2.6 | Feb 11, 2026 |
| Claude Code | 2.0.72 | Claude Opus 4.5 | — | 52.1% ± 2.5 | Nov 24, 2025 |
| OpenHands | 1.1.0 | Claude Opus 4.5 | — | 51.9% ± 2.9 | Nov 24, 2025 |
| OpenCode | — | Claude Opus 4.5 | — | 51.7% | Nov 24, 2025 |
| Terminus 2 | 2.0.0 | Gemini 3 Flash | — | 51.7% ± 3.1 | Dec 17, 2025 |
| Warp | — | Multiple | — | 50.1% ± 2.7 | Nov 11, 2025 |
| Codex CLI | 0.53.0 | GPT-5 | — | 49.6% ± 2.9 | Aug 7, 2025 |
| Terminus 2 | 2.0.0 | GPT-5.1 | — | 47.6% ± 2.8 | Nov 12, 2025 |
| Gemini CLI | — | Gemini 3 Flash | — | 47.4% ± 3.0 | Dec 17, 2025 |
| CAMEL-AI | 1.0 | Claude Sonnet 4.5 | — | 46.5% ± 2.4 | Sep 29, 2025 |
| IndusAGI Coding Agent | — | MiniMax M2.7 | — | 45.1% | Mar 18, 2026 |
| Codex CLI | 0.53.0 | GPT-5-Codex | — | 44.3% ± 2.7 | Sep 15, 2025 |
| OpenHands | 0.60.0 | GPT-5 | — | 43.8% ± 3.0 | Aug 7, 2025 |
| Terminus 2 | 2.0.0 | GPT-5-Codex | — | 43.4% ± 2.9 | Sep 15, 2025 |
| Terminus 2 | 2.0.0 | Kimi K2.5 | — | 43.2% ± 2.9 | Jan 27, 2026 |
| Goose | stable | Claude Sonnet 4.5 | — | 43.1% ± 2.6 | Sep 29, 2025 |
| Crux | — | GPT-5.1-Codex-Mini | — | 43.1% ± 3.0 | Nov 12, 2025 |
| Harness Agent | — | MiniMax M2.7 Highspeed | — | 42.9% ± 2.9 | Mar 18, 2026 |
| Terminus 2 | 2.0.0 | Claude Sonnet 4.5 | — | 42.8% ± 2.8 | Sep 29, 2025 |
| MAYA-V2 | 1.5-official | Claude 4.5 Sonnet | — | 42.7% | Sep 29, 2025 |
| cchuter | — | minimax-m2.5 | — | 42.7% ± 2.8 | Feb 12, 2026 |
| OpenHands | 0.60.0 | Claude Sonnet 4.5 | — | 42.6% ± 2.8 | Sep 29, 2025 |
| Mini-SWE-Agent | — | Claude Sonnet 4.5 | — | 42.5% ± 2.8 | Sep 29, 2025 |
| Terminus 2 | 2.0.0 | Minimax m2.5 | — | 42.2% ± 2.6 | Feb 12, 2026 |
| Mini-SWE-Agent | — | GPT-5-Codex | — | 41.4% ± 2.8 | Sep 15, 2025 |
| Claude Code | 2.0.31 | Claude Sonnet 4.5 | — | 40.1% ± 2.9 | Sep 29, 2025 |
| Terminus 2 | 2.0.0 | DeepSeek-V3.2 | — | 39.5% ± 2.8 | Dec 1, 2025 |
| Terminus 2 | 2.0.0 | Claude Opus 4.1 | — | 38.0% ± 2.6 | Aug 5, 2025 |
| OpenHands | 0.60.0 | Claude Opus 4.1 | — | 36.9% ± 2.7 | Aug 5, 2025 |
| Terminus 2 | 2.0.0 | GPT-5.1-Codex | — | 36.9% ± 3.2 | Nov 12, 2025 |
| Crux | — | MiniMax M2.1 | — | 36.6% ± 2.9 | Dec 23, 2025 |
| Terminus 2 | 2.0.0 | Kimi K2 Thinking | — | 35.7% ± 2.8 | Nov 6, 2025 |
| Goose | stable | Claude Haiku 4.5 | — | 35.5% ± 2.9 | Oct 15, 2025 |
| Terminus 2 | 2.0.0 | GPT-5 | — | 35.2% ± 3.1 | Aug 7, 2025 |
| Mini-SWE-Agent | — | Claude Opus 4.1 | — | 35.1% ± 2.5 | Aug 5, 2025 |
| Claude Code | 2.0.31 | Claude Opus 4.1 | — | 34.8% ± 2.9 | Aug 5, 2025 |
| spoox-o-m | — | GPT-5-Mini | — | 34.8% ± 2.7 | Aug 7, 2025 |
| Mini-SWE-Agent | — | GPT-5 | — | 33.9% ± 2.9 | Aug 7, 2025 |
| Terminus 2 | 2.0.0 | GLM 4.7 | — | 33.4% ± 2.8 | Dec 22, 2025 |
| Crux | — | GLM 4.7 | — | 33.3% ± 2.5 | Dec 22, 2025 |
| Terminus 2 | 2.0.0 | Gemini 2.5 Pro | — | 32.6% ± 3.0 | Jun 17, 2025 |
| Codex CLI | 0.53.0 | GPT-5-Mini | — | 31.9% ± 3.0 | Aug 7, 2025 |
| Terminus 2 | 2.0.0 | MiniMax M2 | — | 30.0% ± 2.7 | Oct 27, 2025 |
| Mini-SWE-Agent | — | Claude Haiku 4.5 | — | 29.8% ± 2.5 | Oct 15, 2025 |
| Terminus 2 | 2.0.0 | MiniMax M2.1 | — | 29.2% ± 2.9 | Dec 23, 2025 |
| OpenHands | 0.60.0 | GPT-5-Mini | — | 29.2% ± 2.8 | Aug 7, 2025 |
| Terminus 2 | 2.0.0 | Claude Haiku 4.5 | — | 28.3% ± 2.9 | Oct 15, 2025 |
| Terminus 2 | 2.0.0 | Kimi K2 Instruct | — | 27.8% ± 2.5 | Jul 11, 2025 |
| Claude Code | 2.0.31 | Claude Haiku 4.5 | — | 27.5% ± 2.8 | Oct 15, 2025 |
| OpenHands | 0.60.0 | Grok 4 | — | 27.2% ± 3.0 | Jul 9, 2025 |
| Dakou Agent | 1.18.72-pre | Qwen 3 Coder 480B | — | 27.2% ± 2.6 | Jul 22, 2025 |
| OpenHands | 0.60.0 | Kimi K2 Instruct | — | 26.7% ± 2.7 | Jul 11, 2025 |
| Mini-SWE-Agent | — | Gemini 2.5 Pro | — | 26.1% ± 2.5 | Jun 17, 2025 |
| Mini-SWE-Agent | — | Grok Code Fast 1 | — | 25.8% ± 2.6 | Aug 28, 2025 |
| Mini-SWE-Agent | — | Grok 4 | — | 25.4% ± 2.9 | Jul 9, 2025 |
| OpenHands | 0.60.0 | Qwen 3 Coder 480B | — | 25.4% ± 2.6 | Jul 22, 2025 |
| little-coder | 0.1.14 | Qwen3.6-35B-A3B | — | 24.6% ± 3.2 | Apr 16, 2026 |
| Terminus 2 | 2.0.0 | GLM 4.6 | — | 24.5% ± 2.4 | Sep 30, 2025 |
| Terminus 2 | 2.0.0 | GPT-5-Mini | — | 24.0% ± 2.5 | Aug 7, 2025 |
| Terminus 2 | 2.0.0 | Qwen 3 Coder 480B | — | 23.9% ± 2.8 | Jul 22, 2025 |
| Terminus 2 | 2.0.0 | Grok 4 | — | 23.1% ± 2.9 | Jul 9, 2025 |
| little-coder | 0.1.13 | Qwen3.6-35B-A3B | — | 23.0% | Apr 16, 2026 |
| Mini-SWE-Agent | — | GPT-5-Mini | — | 22.2% ± 2.6 | Aug 7, 2025 |
| spoox-o-m | — | GPT-5-Nano | — | 21.8% ± 2.8 | Aug 7, 2025 |
| Gemini CLI | 0.11.3 | Gemini 2.5 Pro | — | 19.6% ± 2.9 | Jun 17, 2025 |
| Bash Agent | 1.0.0 | TermiGen-32B | — | 19.3% ± 2.0 | May 14, 2026 |
| Terminus 2 | 2.0.0 | GPT-OSS-120B | — | 18.7% ± 2.7 | Aug 5, 2025 |
| Mini-SWE-Agent | — | Gemini 2.5 Flash | — | 17.1% ± 2.5 | Jun 17, 2025 |
| Terminus 2 | 2.0.0 | AfterQuery-GPT-OSS-20B | — | 17.0% ± 2.5 | Mar 31, 2026 |
| Terminus 2 | 2.0.0 | Gemini 2.5 Flash | — | 16.9% ± 2.4 | Jun 17, 2025 |
| OpenHands | 0.60.0 | Gemini 2.5 Pro | — | 16.4% ± 2.8 | Jun 17, 2025 |
| OpenHands | 0.60.0 | Gemini 2.5 Flash | — | 16.4% ± 2.4 | Jun 17, 2025 |
| Gemini CLI | 0.11.3 | Gemini 2.5 Flash | — | 15.4% ± 2.3 | Jun 17, 2025 |
| Terminus 2 | 2.0.0 | Grok Code Fast 1 | — | 14.2% ± 2.5 | Aug 28, 2025 |
| Mini-SWE-Agent | — | GPT-OSS-120B | — | 14.2% ± 2.3 | Aug 5, 2025 |
| OpenHands | 0.60.0 | Claude Haiku 4.5 | — | 13.9% ± 2.7 | Oct 15, 2025 |
| Codex CLI | 0.53.0 | GPT-5-Nano | — | 11.5% ± 2.3 | Aug 7, 2025 |
| OpenHands | 0.60.0 | GPT-5-Nano | — | 9.9% ± 2.1 | Aug 7, 2025 |
| little-coder | 0.1.24 | Qwen3.5-9B | — | 9.2% ± 2.4 | Mar 2, 2026 |
| Terminus 2 | 2.0.0 | GPT-5-Nano | — | 7.9% ± 1.9 | Aug 7, 2025 |
| Mini-SWE-Agent | — | GPT-5-Nano | — | 7.0% ± 1.9 | Aug 7, 2025 |
| Mini-SWE-Agent | — | GPT-OSS-20B | — | 3.4% ± 1.4 | Aug 5, 2025 |
| Terminus 2 | 2.0.0 | GPT-OSS-20B | — | 3.1% ± 1.5 | Aug 5, 2025 |
Which agents improved most on Terminal-Bench 2.0
Each agent's best run with the earliest model it was listed with, against its best run to date. The gain comes from newer models and from changes to the agent itself; the board does not separate the two.
| Agent | First listed run | Tasks solved | Best run to date | Tasks solved | Gain |
|---|---|---|---|---|---|
| Gemini CLI | Gemini 2.5 Pro · Jun 17, 2025 | 19.6% | Gemini 3.1 Pro · v0.35.0 · Feb 19, 2026 | 61.4% | 41.9 pts |
| spoox-o-m | GPT-5-Mini · Aug 7, 2025 | 34.8% | GPT-5.3-Codex · Feb 5, 2026 | 71.5% | 36.6 pts |
| OpenHands | Gemini 2.5 Pro · Jun 17, 2025 | 16.4% | Claude Opus 4.5 · v1.1.0 · Nov 24, 2025 | 51.9% | 35.5 pts |
| Codex CLI | GPT-5 · Aug 7, 2025 | 49.6% | GPT-5.5 · v0.121.0 · Apr 23, 2026 | 82.2% | 32.6 pts |
| Terminus 2 | Gemini 2.5 Pro · Jun 17, 2025 | 32.6% | GPT-5.3-Codex · v2.0.0 · Feb 5, 2026 | 64.7% | 32.1 pts |
| MAYA-V2 | Claude 4.5 Sonnet · Sep 29, 2025 | 42.7% | Claude 4.6 Opus · v1.5-official · Feb 5, 2026 | 72.1% | 29.4 pts |
| Claude Code | Claude Opus 4.1 · Aug 5, 2025 | 34.8% | Claude Opus 4.6 · v2.1.34 · Feb 5, 2026 | 58.0% | 23.1 pts |
| Mini-SWE-Agent | Gemini 2.5 Pro · Jun 17, 2025 | 26.1% | Claude Sonnet 4.5 · Sep 29, 2025 | 42.5% | 16.5 pts |
| Droid | Gemini 3 Pro · Nov 18, 2025 | 61.1% | GPT-5.3-Codex · Feb 5, 2026 | 77.3% | 16.2 pts |
| Mux | Claude Opus 4.5 · Nov 24, 2025 | 58.4% | GPT-5.3-Codex · Feb 5, 2026 | 74.6% | 16.2 pts |
| CodeBrain-1.5 | Gemini 3 Pro · Nov 18, 2025 | 62.2% | GPT-5.3-Codex · v1.0.0 · Feb 5, 2026 | 75.8% | 13.6 pts |
| SageAgent | Gemini 3 Pro · Nov 18, 2025 | 65.2% | GPT-5.3-Codex · Feb 5, 2026 | 78.4% | 13.3 pts |
| Crux | GPT-5.1-Codex · Nov 12, 2025 | 57.8% | Claude Opus 4.6 · Feb 5, 2026 | 66.9% | 9.1 pts |
| Capy | Claude Opus 4.6 · Feb 5, 2026 | 75.3% | GPT-5.5 · Apr 23, 2026 | 83.2% | 7.9 pts |
| Junie CLI | Gemini 3 Flash · Dec 17, 2025 | 64.3% | Multiple · v507.2 · Mar 7, 2026 | 71.0% | 6.7 pts |
How to read these results
- A row is an agent plus a model. The same agent appears several times with different models and reasoning settings, so the top table shows each agent's best run and the full list sits underneath.
- Each version is its own scale. The task set changes between versions and gets harder, so a score on 4.0 cannot be compared with a score on 2.0. Compare agents within one board.
- Mind the interval. Two runs whose ± ranges overlap are not meaningfully different.
- Cost is for the whole run. It is the total model spend to attempt every task, which makes it a fair way to compare how expensive two set-ups are to operate.
- The date is the model's release date, as listed by the leaderboard, and not the day the run was submitted.
- Coverage is narrow. Only agents that someone submitted are listed. Absence says nothing about quality, and no comparable public benchmark exists yet for support or general-purpose agents.
Source: Terminal-Bench (tbench.ai), hosted by Stanford, Harbor and the Laude Institute, https://www.tbench.ai/; retrieved Oct 6, 2026. Scores belong to their submitters and are shown here for reference with a link to the original.