APEX-Agents厂商自报 · 越高越好
33.5%
成绩出处ARC-AGI v2厂商自报 · 越高越好
77.1%
ARC Prize Verified
成绩出处BrowseComp厂商自报 · 越高越好
85.9%
Search + Python + Browse
成绩出处DeepSWE 1.1来源汇总 · 越高越好
12%
DeepSWE v1.1 leaderboard, mini-swe-agent harness, high effort; Pass@1 12% ± 2%.
成绩出处Finance Agent v2来源汇总 · 越高越好
42.98%
成绩出处FrontierSWE来源汇总 · 越高越好
40%
Gemini CLI
成绩出处GPQA厂商自报 · 越高越好
94.3%
No tools
成绩出处Humanity's Last Exam厂商自报 · 越高越好
51.4%
Search (blocklist) + Code
成绩出处Legal Agent Benchmark来源汇总 · 越高越好
0%
All-pass rate (Harvey held-out set)
成绩出处LiveBench来源汇总 · 越高越好
79.93%
2026-01-08, High
成绩出处LiveCodeBench Pro厂商自报 · 越高越好
2,887分
Elo Rating
成绩出处MCP Atlas厂商自报 · 越高越好
69.2%
成绩出处MMMLU厂商自报 · 越高越好
92.6%
成绩出处MMMU-Pro厂商自报 · 越高越好
80.5%
No tools
成绩出处MRCR v2 (8-needle)厂商自报 · 越高越好
26.3%
1M (pointwise)
成绩出处SciCode厂商自报 · 越高越好
59%
成绩出处SWE-Bench Pro厂商自报 · 越高越好
54.2%
Single attempt
成绩出处SWE-Bench Verified厂商自报 · 越高越好
80.6%
Single attempt
成绩出处t2-bench厂商自报 · 越高越好
99.3%
Telecom
成绩出处Terminal-Bench 2.0厂商自报 · 越高越好
68.5%
Terminus-2 harness
成绩出处