Narev
  • Twitter
  • Create

  • Evals
    • Models
    • Hardware

  • Catalog

> Hub > Evals > Models > …

loading run

> Hub > Evals > Models > Frontier Multi-Domain Advanced Reasoning Benchmark

Evaluates advanced problem-solving, formal logic, mathematical reasoning, theoretical computer science, and epistemological edge cases designed to test frontier language models.

Scorer: Factuality

loading chart
Score vs cost per task
#scoremodel
cost
price
ttfttotal tok
1100%openai/gpt-5.6-luna-pro
$0.00036
$0.10 / $0.60 per 1M
3.41s2,309.9
2100%qwen/qwen3.8-max
$0.00252
$2.00 / $6.00 per 1M
1.59s514
3failedanthropic/claude-opus-5-fast
—
$10.00 / $50.00 per 1M
—0

Prompts

  • A standard fair 6-sided die is rolled repeatedly until a 6 appears. What is the expected number of t…
  • According to the No-Cloning Theorem in quantum mechanics, which of the following operations is funda…
  • Consider the Monty Hall problem with 4 doors, where behind 1 door is a car and behind 3 are goats. Y…
  • In computational complexity theory, which of the following relations is currently proven to be true?…
  • In distributed systems, according to the CAP Theorem formulated by Eric Brewer, which three properti…
  • In formal logic, if a system of first-order arithmetic is consistent, which of the following stateme…
  • In propositional logic, which of the following is logically equivalent to the conditional statement …
  • In standard Zermelo-Fraenkel set theory with the Axiom of Choice (ZFC), which of the following propo…
  • What is the general term for the linguistic phenomenon where a sentence has a structure that initial…
  • What is the value of the infinite sum S = 1/2 + 1/4 + 1/8 + 1/16 + ... ? A) 1 B) 2 C) 1/2 D) Diverge…