Evaluates language models on core Polish grammar, orthography, morphology, syntax, and idiomatic expressions.
Scorer: Factuality
| # | score | model | cost price | ttft | total tok |
|---|---|---|---|---|---|
| 1 | 100% | gpt-5.6-luna | $0.00003 $0.20 / $1.20 per 1M | 1.18s | 97.2 |
| 2 | 100% | anthropic/claude-opus-4.8 | $0.00064 $5.00 / $25.00 per 1M | 1.27s | 116.8 |
| 3 | 100% | z-ai/glm-5.2 | $0.00129 $0.60 / $2.00 per 1M | 8.44s | 409.1 |
Prompts