Evaluates advanced problem-solving, formal logic, mathematical reasoning, theoretical computer science, and epistemological edge cases designed to test frontier language models.
Scorer: Factuality
| # | score | model | cost price | ttft | total tok |
|---|---|---|---|---|---|
| 1 | 100% | openai/gpt-5.6-luna-pro | $0.00036 $0.10 / $0.60 per 1M | 3.41s | 2,309.9 |
| 2 | 100% | qwen/qwen3.8-max | $0.00252 $2.00 / $6.00 per 1M | 1.59s | 514 |
| 3 | failed | anthropic/claude-opus-5-fast | — $10.00 / $50.00 per 1M | — | 0 |
Prompts