正在读取模型档案与评测数据…
正在读取模型档案与评测数据…
Shanghai AI Laboratory·1 个数据来源·更新于 2026-09-26
只在同一来源、同一口径下比较。切换来源可查看不同测试结果。
当前配置:各基准设置见下表
综合分数和同榜图表来自 LLM Stats 榜单(采集于 2026-09-25);下方逐项成绩来自模型页。榜单出处
LLM Stats Score · 同一来源与榜单口径,分数越高越好
| 模型与配置 | LLM Stats Score | 每任务成本 |
|---|---|---|
| Muse Spark 1.3 | 53.8 | 未公布 |
| Kimi K3 | 52.5 | 未公布 |
| GLM-5.3 | 52.1 | 未公布 |
| Atria Dawn Preview · 当前模型 | 51.6 | 未公布 |
| Qwen3.8 Max | 51.5 | 未公布 |
| DeepSeek-V4.1-Flash | 51.4 | 未公布 |
| DeepSeek-V4-Pro-0813 | 51.1 | 未公布 |
不同来源的分数不能直接相加或横向比较;评测任务成本不是 API Token 单价。
12 项已收录成绩 · 保留原始单位和测试条件
AutomationBench. Official Atria Dawn Preview Hugging Face model card evaluation table.
成绩出处BFCL v4. Official Atria Dawn Preview Hugging Face model card evaluation table.
成绩出处BrowseComp. Official Atria Dawn Preview Hugging Face model card evaluation table.
成绩出处CyberGym. Official Atria Dawn Preview Hugging Face model card evaluation table.
成绩出处DeepSearchQA. Official Atria Dawn Preview Hugging Face model card evaluation table.
成绩出处GDPval row (Elo-style scale, e.g. Fable 5.1 reports 1853 on the same scale). Mapped to gdpval-aa (max_score 3000) rather than the fractional gdpval benchmark because the reported value is on the Elo index scale, not a 0-1 win rate. Official Atria Dawn Preview Hugging Face model card evaluation table.
成绩出处MLE-bench Lite. Official Atria Dawn Preview Hugging Face model card evaluation table.
成绩出处SkillsBench. Official Atria Dawn Preview Hugging Face model card evaluation table.
成绩出处SWE-bench Pro. Official Atria Dawn Preview Hugging Face model card evaluation table.
成绩出处τ³-Bench Banking. Official Atria Dawn Preview Hugging Face model card evaluation table.
成绩出处Terminal-Bench 2.1. Official Atria Dawn Preview Hugging Face model card evaluation table.
成绩出处WideSearch. Official Atria Dawn Preview Hugging Face model card evaluation table.
成绩出处只展示有可靠来源的资料;优先采用厂商官网与官方模型仓库。
美元 / 百万 Token。不同服务商或测试配置的报价分别列出。
同名模型通过开发机构和明确版本关联。预览版、日期版本和不同规模模型分别建档。
公开模型页与公开榜单;自报成绩单独标识