Evaluates core Python programming knowledge across language semantics, standard library behavior, scoping, memory model, and object-oriented paradigms.
Scorer: Factuality
| # | score | model | cost price | ttft | total tok |
|---|---|---|---|---|---|
| 1 | 100% | gpt-5.3-codex | $0.00023 $1.75 / $14.00 per 1M | 1.07s | 98.8 |
| 2 | 100% | moonshotai/kimi-k2.7-code | $0.00045 $0.69 / $3.49 per 1M | 0.72s | 203.6 |
| 3 | 90% | qwen/qwen3-coder-next | $0.00001 $0.12 / $0.80 per 1M | 0.90s | 100.1 |
Prompts