本文へ移動

AIモデル比較

Terminal-Bench 4.0

Terminal-Benchは、AIエージェントがターミナルを使って作業を完了する能力を測るベンチマークです。4.0の解決率を0〜100%で示し、数値が高いほど多くの作業を完了できたことを表します。
モデル・実行構成%
Claude Sonnet 5.5max 70.60
Claude Sonnet 5.5xhigh 61.50
GPT-6 Astramax 58.1895% CI half-width ±2.79 percentage points
Fable 5.1max 57.8895% CI half-width ±3.76 percentage points
GPT-6 Astraxhigh 57.8895% CI half-width ±2.72 percentage points
GPT-6 Astrahigh 57.8895% CI half-width ±2.97 percentage points
Fable 5.1xhigh 57.8895% CI half-width ±3.36 percentage points
Fable 5.1high 54.5595% CI half-width ±3.44 percentage points
GPT-6 Astramedium 54.2495% CI half-width ±2.66 percentage points
Fable 5.1medium 53.9495% CI half-width ±3.39 percentage points
Opus 5max 51.8295% CI half-width ±3.39 percentage points
GPT-6 Astralow 50.6195% CI half-width ±2.75 percentage points
Fable 5max 44.5595% CI half-width ±3.85 percentage points
Fable 5.1low 43.3395% CI half-width ±3.61 percentage points
Claude Sonnet 5.5high 43.00
GLM-5.3max 41.8295% CI half-width ±3.23 percentage points
GPT-5.6 Solmax 37.2795% CI half-width ±3.78 percentage points
Claude Sonnet 5.5medium 28.80
Opus 4.8max 23.6495% CI half-width ±3.56 percentage points
GPT-5.6 Terramax 21.5295% CI half-width ±3.25 percentage points
Grok 4.6high 20.3095% CI half-width ±3.09 percentage points
Claude Sonnet 5.5low 20.00
Gemini 3.8 Flashhigh 19.0995% CI half-width ±3.36 percentage points
GPT-5.6 Lunamax 17.2795% CI half-width ±2.85 percentage points
Grok 4.5high 12.4295% CI half-width ±2.62 percentage points
Sonnet 5max 12.4295% CI half-width ±3.06 percentage points
Gemini 3.7 Flashhigh 11.2195% CI half-width ±2.45 percentage points