Agents' Last Exam厂商自报 · 越高越好
51.2%
Score
成绩出处AndroidWorld厂商自报 · 越高越好
84.5%
成绩出处CharXiv-R厂商自报 · 越高越好
90.6%
With code interpreter
成绩出处ClawEval-MM厂商自报 · 越高越好
60.4%
Average score across three trials
成绩出处CoWorkBench厂商自报 · 越高越好
73.9%
成绩出处DeepSWE 1.1厂商自报 · 越高越好
58.7%
Best of Claude Code and mini-SWE-agent, temp=1.0, top_p=0.95, 256K context
成绩出处GPQA厂商自报 · 越高越好
91.7%
Diamond split
成绩出处Humanity's Last Exam厂商自报 · 越高越好
35.9%
Judged by GPT-4o
成绩出处IFBench厂商自报 · 越高越好
81.3%
成绩出处Job Bench厂商自报 · 越高越好
55.7%
成绩出处LiveCodeBench v6厂商自报 · 越高越好
91.9%
成绩出处LVBench厂商自报 · 越高越好
76.6%
成绩出处MathVision厂商自报 · 越高越好
95.7%
With code interpreter
成绩出处NL2Repo厂商自报 · 越高越好
48.1%
Claude Code harness
成绩出处OSWorld 2.0厂商自报 · 越高越好
19.4%
Binary completion
成绩出处RealWorldQA厂商自报 · 越高越好
88.5%
成绩出处RecreationBench厂商自报 · 越高越好
49.9%
成绩出处SWE-bench Multilingual厂商自报 · 越高越好
81%
mini-SWE-agent, temp=1.0, top_p=0.95, 256K context
成绩出处SWE-Bench Pro厂商自报 · 越高越好
62.5%
Claude Code, temp=1.0, top_p=0.95, 256K context
成绩出处Toolathlon厂商自报 · 越高越好
73.5%
Verified split; Pass@1
成绩出处Vision2Web厂商自报 · 越高越好
64%
Claude Code harness, judged by gpt-5.4-2026-03-05
成绩出处