AA-Briefcase厂商自报 · 越高越好
917分
Elo score; Artificial Analysis (max_score 3000)
成绩出处AIME 2026厂商自报 · 越高越好
95.5%
成绩出处ARC-AGI厂商自报 · 越高越好
84%
ARC-AGI-1
成绩出处ARC-AGI v2厂商自报 · 越高越好
40.1%
ARC-AGI-2
成绩出处Artificial Analysis厂商自报 · 越高越好
40%
Artificial Analysis Intelligence Index v4.1. Score 40 (vendor model card).
成绩出处BrowseComp厂商自报 · 越高越好
77.4%
With context management
成绩出处CharXiv-R厂商自报 · 越高越好
77.4%
CharXiv RQ original
成绩出处CritPT厂商自报 · 越高越好
8.3%
成绩出处GDPval-AA厂商自报 · 越高越好
1,269分
GDPval-AA v2 Elo (max_score 3000)
成绩出处Global-MMLU-Lite厂商自报 · 越高越好
86.7%
成绩出处GPQA厂商自报 · 越高越好
89.5%
GPQA Diamond
成绩出处Humanity's Last Exam厂商自报 · 越高越好
31.6%
Text only
成绩出处IFBench厂商自报 · 越高越好
82.2%
成绩出处MCP Atlas厂商自报 · 越高越好
79.6%
Public set
成绩出处MMMU-Pro厂商自报 · 越高越好
74%
MMMU Pro Standard 10
成绩出处SciCode厂商自报 · 越高越好
48.7%
成绩出处SimpleQA Verified厂商自报 · 越高越好
20.6%
成绩出处SWE-Bench Pro厂商自报 · 越高越好
55.9%
SWE-Bench Pro public
成绩出处SWE-Bench Verified厂商自报 · 越高越好
80.2%
Bash-only harness (same as Inkling 77.6%)
成绩出处Tau3 Banking厂商自报 · 越高越好
15.5%
成绩出处Terminal-Bench 2.1厂商自报 · 越高越好
64.7%
Vendor best harness (model card)
成绩出处Toolathlon厂商自报 · 越高越好
54.4%
Toolathlon Verified
成绩出处VoiceBench Avg厂商自报 · 越高越好
90.1%
成绩出处