Narev
  • Twitter
  • Create

  • Evals
    • Models
    • Hardware

  • Catalog

> Hub > Evals > Models > …

loading run

> Hub > Evals > Models > Advanced Multidisciplinary Frontier Reasoning Benchmark

A high-difficulty benchmark testing frontier reasoning capabilities across mathematical logic, formal semantics, computability theory, quantum physics, game theory, and counter-intuitive probability.

Scorer: Factuality

loading chart
Score vs cost per task
#scoremodel
cost
price
ttfttotal tok
1100%qwen/qwen3.8-max
$0.01824
$2.00 / $6.00 per 1M
1.24s3,151.8
290%openai/gpt-5.6-luna-pro
$0.00089
$0.10 / $0.60 per 1M
8.14s3,783.8
372%anthropic/claude-opus-4.8
$0.00487
$5.00 / $25.00 per 1M
1.39s341.2

Prompts

  • Consider a standard three-player game in extensive form with perfect information and no chance moves…
  • Consider an ideal gas undergoing a reversible Otto cycle composed of four steps. Which sequence accu…
  • In abstract algebra, let G be a finite group of order 105 (3 × 5 × 7). Applying the Sylow theorems, …
  • In complexity theory, assuming the standard complexity hierarchy does not collapse, which of the fol…
  • In computability theory, let K = {x | φ_x(x) halts} be the standard halting problem set. Which of th…
  • In distributed computing under the asynchronous message-passing model without failure detectors, the…
  • In formal semantics and pragmatics, which of the following sentences exhibits a genuine presuppositi…
  • In quantum mechanics, according to the Wigner-Eckart theorem, the matrix element of an irreducible t…
  • In standard Zermelo-Fraenkel set theory with the Axiom of Choice (ZFC), which of the following state…
  • Three fair, standard 6-sided dice are rolled simultaneously. What is the conditional probability tha…