A high-difficulty benchmark testing frontier reasoning capabilities across mathematical logic, formal semantics, computability theory, quantum physics, game theory, and counter-intuitive probability.
Scorer: Factuality
| # | score | model | cost price | ttft | total tok |
|---|---|---|---|---|---|
| 1 | 100% | qwen/qwen3.8-max | $0.01824 $2.00 / $6.00 per 1M | 1.24s | 3,151.8 |
| 2 | 90% | openai/gpt-5.6-luna-pro | $0.00089 $0.10 / $0.60 per 1M | 8.14s | 3,783.8 |
| 3 | 72% | anthropic/claude-opus-4.8 | $0.00487 $5.00 / $25.00 per 1M | 1.39s | 341.2 |
Prompts