Evaluates whether an assistant recommends an appropriately small, fast, and cost-effective model—or deterministic code—for elementary arithmetic, while accounting for accuracy, scale, and operational constraints.
Scorer: ClosedQA
| # | score | model | cost price | ttft | total tok |
|---|---|---|---|---|---|
| 1 | 70% | gemini-2.5-flash-lite | $0.00048 $0.10 / $0.40 per 1M | 0.43s | 1,240.4 |
| 2 | 70% | gpt-5-nano | $0.00119 $0.05 / $0.40 per 1M | 18.52s | 3,017.2 |
| 3 | 50% | amazon/nova-micro-v1 | $0.00007 $0.04 / $0.14 per 1M | 0.52s | 565.6 |