Evaluates a model's accuracy on ten plain-number arithmetic problems spanning common operations, signs, decimals, percentages, fractions, order of operations, and short word problems.
Scorer: NumericDiff
| # | score | model | cost price | ttft | total tok |
|---|---|---|---|---|---|
| 1 | 100% | gemini-3.5-flash-lite | $0.00001 $0.30 / $2.50 per 1M | 0.56s | 25.4 |
| 2 | 99% | microsoft/phi-4 | $0.00000 $0.07 / $0.14 per 1M | 0.32s | 29.3 |
| 3 | 98% | gpt-5.4-nano | $0.00001 $0.20 / $1.25 per 1M | 0.37s | 31.3 |