AdvancedIF厂商自报 · 越高越好
85%
Rubric-level
成绩出处AIME 2025厂商自报 · 越高越好
97%
成绩出处AIME 2026厂商自报 · 越高越好
94.5%
成绩出处AIR-Bench厂商自报 · 越高越好
88%
成绩出处BFCL-v3厂商自报 · 越高越好
72%
成绩出处CorpusQA厂商自报 · 越高越好
82%
成绩出处CyberSecEval 4厂商自报 · 越高越好
63%
Insecure code: Autocomplete
成绩出处GPQA厂商自报 · 越高越好
84.2%
Diamond
成绩出处GraphWalks厂商自报 · 越高越好
90%
F1, ≤128k context
成绩出处HealthBench Professional厂商自报 · 越高越好
35%
成绩出处HMMT Feb 26厂商自报 · 越高越好
84.9%
成绩出处IFBench厂商自报 · 越高越好
69%
成绩出处LiveCodeBench v6厂商自报 · 越高越好
87.7%
成绩出处LongBench v2厂商自报 · 越高越好
61%
256k context
成绩出处LongFact厂商自报 · 越高越好
98%
成绩出处MedXpertQA厂商自报 · 越高越好
43%
Text
成绩出处MMLU-Pro厂商自报 · 越高越好
85%
成绩出处Multi-Challenge厂商自报 · 越高越好
53%
成绩出处SimpleQA Verified厂商自报 · 越高越好
31%
成绩出处SWE-Bench Pro厂商自报 · 越高越好
52.8%
256k context
成绩出处SWE-Bench Verified厂商自报 · 越高越好
73.5%
256k context
成绩出处Terminal-Bench 2.0厂商自报 · 越高越好
46%
256k context
成绩出处TruthfulQA厂商自报 · 越高越好
88%
成绩出处