AIモデル比較
Terminal-Bench 4.0
Terminal-Benchは、AIエージェントがターミナルを使って作業を完了する能力を測るベンチマークです。4.0の解決率を0〜100%で示し、数値が高いほど多くの作業を完了できたことを表します。
| モデル・実行構成 | % |
|---|---|
| Claude Sonnet 5.5max | 70.60 |
| Claude Sonnet 5.5xhigh | 61.50 |
| GPT-6 Astramax | 58.1895% CI half-width ±2.79 percentage points |
| Fable 5.1max | 57.8895% CI half-width ±3.76 percentage points |
| GPT-6 Astraxhigh | 57.8895% CI half-width ±2.72 percentage points |
| GPT-6 Astrahigh | 57.8895% CI half-width ±2.97 percentage points |
| Fable 5.1xhigh | 57.8895% CI half-width ±3.36 percentage points |
| Fable 5.1high | 54.5595% CI half-width ±3.44 percentage points |
| GPT-6 Astramedium | 54.2495% CI half-width ±2.66 percentage points |
| Fable 5.1medium | 53.9495% CI half-width ±3.39 percentage points |
| Opus 5max | 51.8295% CI half-width ±3.39 percentage points |
| GPT-6 Astralow | 50.6195% CI half-width ±2.75 percentage points |
| Fable 5max | 44.5595% CI half-width ±3.85 percentage points |
| Fable 5.1low | 43.3395% CI half-width ±3.61 percentage points |
| Claude Sonnet 5.5high | 43.00 |
| GLM-5.3max | 41.8295% CI half-width ±3.23 percentage points |
| GPT-5.6 Solmax | 37.2795% CI half-width ±3.78 percentage points |
| Claude Sonnet 5.5medium | 28.80 |
| Opus 4.8max | 23.6495% CI half-width ±3.56 percentage points |
| GPT-5.6 Terramax | 21.5295% CI half-width ±3.25 percentage points |
| Grok 4.6high | 20.3095% CI half-width ±3.09 percentage points |
| Claude Sonnet 5.5low | 20.00 |
| Gemini 3.8 Flashhigh | 19.0995% CI half-width ±3.36 percentage points |
| GPT-5.6 Lunamax | 17.2795% CI half-width ±2.85 percentage points |
| Grok 4.5high | 12.4295% CI half-width ±2.62 percentage points |
| Sonnet 5max | 12.4295% CI half-width ±3.06 percentage points |
| Gemini 3.7 Flashhigh | 11.2195% CI half-width ±2.45 percentage points |
