正在读取模型档案与评测数据…
正在读取模型档案与评测数据…
Anthropic·4 个数据来源·更新于 2026-09-26
只在同一来源、同一口径下比较。切换来源可查看不同测试结果。
当前配置:各基准设置见下表
综合分数和同榜图表来自 LLM Stats 榜单(采集于 2026-09-25);下方逐项成绩来自模型页。榜单出处
LLM Stats Score · 同一来源与榜单口径,分数越高越好
| 模型与配置 | LLM Stats Score | 每任务成本 |
|---|---|---|
| Claude Opus 5.5 · 当前模型 | 59.7 | 未公布 |
| GPT-6 Astra | 59.5 | 未公布 |
| Claude Opus 5 | 54.8 | 未公布 |
| Claude Fable 5.1 | 54.5 | 未公布 |
| GPT-5.6 Sol | 54.4 | 未公布 |
| Claude Mythos Preview | 54.3 | 未公布 |
| Claude Fable 5 | 54 | 未公布 |
不同来源的分数不能直接相加或横向比较;评测任务成本不是 API Token 单价。
31 项已收录成绩 · 保留原始单位和测试条件
System Card Table 8.1.A. AA-Briefcase v1.1 Elo at max effort (max_score 3000).
成绩出处System Card §8.9. ArXivMath August 2026 (57 problems); without tools; max effort; averaged over four attempts per problem. Evaluated internally without safeguards classifiers.
成绩出处Opus 5.5 with production safeguards enabled. Zapier AutomationBench; runs without fallback models, so safeguard interventions counted as failures. Adaptive thinking at max effort.
成绩出处System Card §8.13.2. BenchCAD Vision2Code 1k subset; voxel IoU without tools; adaptive thinking at max effort; averaged over five runs.
成绩出处System Card §8.13.2. BenchCAD Vision2Code 1k subset; voxel IoU with tools; adaptive thinking at max effort; averaged over five runs.
成绩出处System Card §8.17.1. BioMysteryBench Human Solvable subset 89.3%; adaptive thinking at max effort; bash/file tools with package allow-list.
成绩出处Opus 5.5 with production safeguards enabled. Chartography with tools; adaptive thinking at max effort.
成绩出处Opus 5.5 with production safeguards enabled. CursorBench 4.0; adaptive thinking at max effort.
成绩出处System Card §8.3. DeepSWE v1.1; 5-trial average. Adaptive thinking at max effort.
成绩出处Opus 5.5 with production safeguards enabled. FrontierCode v1.1 Main subset; adaptive thinking at max effort.
成绩出处System Card §8.7. FrontierSWE v2; Proximal harness; max effort; 5 trials per task; mean across trials.
成绩出处Opus 5.5 with production safeguards enabled. GDPval-AA v2.1 Elo (max_score 3000); adaptive thinking at max effort.
成绩出处System Card §8.16.1. GMMLU average accuracy 94.3% across 42 languages; adaptive thinking at max effort; single trial; no tools.
成绩出处System Card §8.15.1. HealthBench length-adjusted score 60.6% (penalizes verbose responses); adaptive thinking at max effort; 5-trial average. Prefer length-adjusted over raw 68.1% to match other catalog models and avoid duplicate benchmark_id records.
成绩出处System Card Table 8.1.A length-adjusted score 65.6% (prefer Table 8.1.A over §8.15.2 raw 77.1%). Adaptive thinking at max effort; 5-trial average; safety classifiers enabled with refusal fallback to Claude Opus 5. Length-adjusted per HealthBench Professional paper method.
成绩出处System Card Table 8.1.A / §8.11.1. HLE no tools (reasoning-only); adaptive thinking at max effort. Alongside with-tools 67.7% already cataloged from the announcement table.
成绩出处Opus 5.5 with production safeguards enabled. Humanity's Last Exam with tools; adaptive thinking at max effort.
成绩出处System Card §8.17.2. LatchBio SingleCellBench 61.2%; adaptive thinking at max effort; bash/file tools.
成绩出处System Card §8.17.2. LatchBio SpatialBench Verified 72.0%; adaptive thinking at max effort; bash/file tools.
成绩出处System Card §8.14.2. Legal Agent Benchmark held-out set; all-pass rate 8.3% at max effort (AA harness). Mean criterion-pass rate 91.2% noted in card but not stored as a separate catalog metric (all-pass is the standard LAB reporting id).
成绩出处System Card §8.16.2. MILU average accuracy 93.1% across 11 languages; adaptive thinking at max effort; averaged over 5 trials; no tools.
成绩出处System Card §8.14.1. OfficeQA full set; max effort; mean of five runs; agentic harness with extracted-text documents and code-execution tools.
成绩出处System Card §8.14.1. OfficeQA Pro (133-question subset); max effort; mean of five runs.
成绩出处Opus 5.5 with production safeguards enabled. OSWorld 2.0 partial score (announcement reports partial, not strict); adaptive thinking at max effort.
成绩出处System Card §8.10.1. ProgramBench 166 golden tasks; mini-swe-agent harness; no 6-hour time limit; hidden test pass rate against tests the reference binary passes.
成绩出处System Card Table 8.1.A / §8.2. Adaptive thinking at max effort; 5-trial average. 300 problems across nine programming languages.
成绩出处System Card Table 8.1.A / §8.2. Adaptive thinking at max effort; 5-trial average. Visual context (screenshots and design mockups).
成绩出处System Card Table 8.1.A / §8.2. Adaptive thinking at max effort; 5-trial average. Production safeguards as in card capability summary.
成绩出处Opus 5.5 with production safeguards enabled. Terminal-Bench 4.0 accuracy 66.4% at xhigh effort (Claude Code harness). Unless otherwise noted, other Opus 5.5 results on this announcement table use adaptive thinking at max effort; TB4 is explicitly reported at xhigh.
成绩出处Opus 5.5 with production safeguards enabled. Terminal-Bench-Science 0.1; adaptive thinking at max effort.
成绩出处System Card §8.14.5. Toolathlon Verified Pass@1 77.8% (3-trial average across 108 tasks). Adaptive thinking at max effort; safety/sandbox stops counted as failures.
成绩出处只展示有可靠来源的资料;优先采用厂商官网与官方模型仓库。
美元 / 百万 Token。不同服务商或测试配置的报价分别列出。
| 来源 / 服务商 | 输入 | 输出 | 缓存读取 |
|---|---|---|---|
| Anthropic | $4 | $20 | $0.2 |
同名模型通过开发机构和明确版本关联。预览版、日期版本和不同规模模型分别建档。
公开评测数据;保留来源原始口径
公开模型页与公开榜单;自报成绩单独标识
公开评测数据;保留来源原始口径
公开模型目录与 API 报价
| 来源 / 配置 | 采集时间 | 原始数据 |
|---|---|---|
| Artificial AnalysisClaude Opus 5.5 (max with fallback)SHA-256 b9324ff6388c2fd3… | 2026-09-25T19:40:27Z | 查看来源 |
| Artificial AnalysisClaude Opus 5.5 (xhigh with fallback)SHA-256 b9324ff6388c2fd3… | 2026-09-25T19:40:27Z | 查看来源 |
| Artificial AnalysisClaude Opus 5.5 (high with fallback)SHA-256 b9324ff6388c2fd3… | 2026-09-25T19:40:27Z | 查看来源 |
| Artificial AnalysisClaude Opus 5.5 (medium with fallback)SHA-256 b9324ff6388c2fd3… | 2026-09-25T19:40:27Z | 查看来源 |
| Artificial AnalysisClaude Opus 5.5 (low with fallback)SHA-256 b9324ff6388c2fd3… | 2026-09-25T19:40:27Z | 查看来源 |
| Artificial AnalysisClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)SHA-256 5beb2f0dcdc468c1… | 2026-09-26T00:27:47Z | 查看来源 |
| Artificial AnalysisClaude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)SHA-256 5beb2f0dcdc468c1… | 2026-09-26T00:27:47Z | 查看来源 |
| Artificial AnalysisClaude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback)SHA-256 5beb2f0dcdc468c1… | 2026-09-26T00:27:47Z | 查看来源 |
| Artificial AnalysisClaude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)SHA-256 5beb2f0dcdc468c1… | 2026-09-26T00:27:47Z | 查看来源 |
| Artificial AnalysisClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)SHA-256 5beb2f0dcdc468c1… | 2026-09-26T00:27:47Z | 查看来源 |
| LLM StatsClaude Opus 5.5 | 2026-09-25T19:40:20Z | 查看来源 |
| LLM StatsClaude Opus 5.5SHA-256 26071d96a7e4b355… | 2026-09-26T00:28:55Z | 查看来源 |
| BenchLMClaude Opus 5.5 | 2026-09-24 | 查看来源 |
| OpenRouterClaude Opus 5.5SHA-256 4c26b1452fbd2acd… | 2026-09-26T00:28:52Z | 查看来源 |