Leaderboard / anthropic/claude-opus-4-8
claude-opus-4-8
anthropic · Tasks passed 91.8% · Finance Index 91.7
Tasks passed
91.8%
Overall accuracy
90.9%
Finance Index
91.7
Consistency
2.7/3 avg
Est. eval cost
$1.23
Median latency
4.2s
Median output speed
18 tok/s
- Harness version
- 0.2.0
- Task set
- v2
- Runs per task
- 3
- Evaluated
- 7/13/2026, 2:27:51 PM
- Run ID
- 8ac810ac-76f0-4fd8-88cd-718cb1f3b150
- Results file
- Download JSON
Pass@1 by Difficulty
Easy
—
Medium
—
Hard
—
Scores by domain
Tasks passed (%) · top 1 models
Quant
Per-attempt drill-down requires a local result JSON in `results/`. Aggregate scores are shown from the published leaderboard.