正在读取模型档案与评测数据…
正在读取模型档案与评测数据…
xai·2 个数据来源·更新于 2026-09-26
只在同一来源、同一口径下比较。切换来源可查看不同测试结果。
当前配置:各基准设置见下表
不同来源的分数不能直接相加或横向比较;评测任务成本不是 API Token 单价。
11 项已收录成绩 · 保留原始单位和测试条件
BioLP-Bench accuracy - Model-graded evaluation measuring ability to find and correct mistakes in common biological laboratory protocols.
成绩出处CloningScenarios accuracy - Expert-level multi-step reasoning questions about difficult genetic cloning scenarios in multiple-choice format.
成绩出处Creative Writing v3 - Normalized Elo score (Elo/2000). Evaluates creative writing capabilities across 32 writing prompts with 3 iterations per prompt.
成绩出处CyBench success rate - Suite of Capture-the-Flag (CTF) challenges measuring agentic cyber attack capabilities. Unguided success rate where agents complete tasks end-to-end without guidance.
成绩出处EQ-Bench Emotional Intelligence Benchmark - Normalized Elo score. Evaluates active emotional intelligence abilities, understanding, insight, empathy, and interpersonal skills.
成绩出处FigQA accuracy - Multiple-choice benchmark on interpreting scientific figures from biology papers. Multimodal reasoning evaluation.
成绩出处Grok 4.1 Thinking (code name: quasarflux) - LMArena Text Leaderboard Elo rating. Holds #1 overall position with 1483 Elo, a commanding margin of 31 points over the highest non-xAI model.
成绩出处MASK honesty score (1 - dishonesty rate) - Measures whether models faithfully report their beliefs when pressured to lie. Dishonesty rate: 0.49. Lower dishonesty rates indicate better honesty.
成绩出处ProtocolQA accuracy - Multiple-choice benchmark on troubleshooting failed experimental outcomes from common biological laboratory protocols.
成绩出处Virology Capabilities Test (VCT) accuracy - Expert-level multiple-choice benchmark measuring capability to troubleshoot complex virology laboratory protocols. Text-only questions evaluated.
成绩出处WMDP Cybersecurity accuracy - Multiple-choice benchmark on dual-use cyber knowledge. Evaluates model's capacity to enable malicious actors to engage in offensive cyber operations.
成绩出处只展示有可靠来源的资料;优先采用厂商官网与官方模型仓库。
美元 / 百万 Token。不同服务商或测试配置的报价分别列出。
同名模型通过开发机构和明确版本关联。预览版、日期版本和不同规模模型分别建档。