Evaluates whether an assistant provides practical, well-reasoned model recommendations for Finnish and other language-specific use cases, accounting for task, quality, deployment, and cost constraints.
Scorer: Possible
| # | score | model | cost price | ttft | total tok |
|---|---|---|---|---|---|
| 1 | 100% | mistral-large-2512 | $0.00239 $0.50 / $1.50 per 1M | 0.70s | 1,610.1 |
| 2 | 100% | gpt-5.5 | $0.02995 $5.00 / $30.00 per 1M | 0.43s | 1,018.6 |
| 3 | running | anthropic/claude-opus-5 | — $5.00 / $25.00 per 1M | — | 0 |
Prompts