Evaluates a model's ability to solve elementary arithmetic and basic word problems while returning only the numerical answer.
Scorer: NumericDiff
| # | score | model | cost price | ttft | total tok |
|---|---|---|---|---|---|
| 1 | 100% | qwen/qwen3.5-9b | $0.00005 $0.10 / $0.15 per 1M | 0.71s | 361.1 |
| 2 | 100% | gpt-5.5 | $0.00078 $5.00 / $30.00 per 1M | 0.32s | 46 |
| 3 | running | anthropic/claude-opus-5 | — $5.00 / $25.00 per 1M | — | 0 |