API-Bank厂商自报 · 越高越好
92%
0-shot
成绩出处ARC-C厂商自报 · 越高越好
96.9%
0-shot
成绩出处BFCL厂商自报 · 越高越好
88.5%
0-shot
成绩出处DROP厂商自报 · 越高越好
84.8%
0-shot
成绩出处Gorilla Benchmark API Bench厂商自报 · 越高越好
35.3%
0-shot
成绩出处GPQA厂商自报 · 越高越好
50.7%
0-shot
成绩出处GSM8k厂商自报 · 越高越好
96.8%
8-shot, CoT, em_maj1@1
成绩出处HumanEval厂商自报 · 越高越好
89%
0-shot, pass@1
成绩出处IFEval厂商自报 · 越高越好
88.6%
Standard
成绩出处MATH厂商自报 · 越高越好
73.8%
0-shot, CoT, final_em
成绩出处MBPP EvalPlus厂商自报 · 越高越好
88.6%
0-shot, base, pass@1
成绩出处MMLU厂商自报 · 越高越好
87.3%
5-shot, macro_avg/acc
成绩出处MMLU (CoT)厂商自报 · 越高越好
88.6%
0-shot, macro_avg/acc
成绩出处MMLU-Pro厂商自报 · 越高越好
73.3%
5-shot, CoT, micro_avg/acc_char
成绩出处Multilingual MGSM (CoT)厂商自报 · 越高越好
91.6%
0-shot, CoT, em
成绩出处Multipl-E HumanEval厂商自报 · 越高越好
75.2%
0-shot, pass@1
成绩出处Multipl-E MBPP厂商自报 · 越高越好
65.7%
0-shot, pass@1
成绩出处Nexus厂商自报 · 越高越好
58.7%
0-shot, macro_avg/acc
成绩出处