KINA Benchmark

A high-density benchmark spanning 261 disciplines with game-theoretic annotation

KINA Benchmark

This page describes KINA (Knowledge Index of Noah’s Ark), a benchmark for evaluating the disciplinary knowledge boundaries of frontier AI models, built at 2077AI.

Existing benchmarks for evaluating AI knowledge suffer from three structural problems: they prioritize scale or extreme difficulty over disciplinary representativeness, they are vulnerable to data contamination, and they rely on blind-trust annotations prone to “lazy consensus” — where annotators agree without independent reasoning. KINA addresses all three.

KINA spans 261 disciplines across the full breadth of human knowledge, organized into a three-level taxonomy. Each question uses a 10-option format to suppress guessing, with a strict difficulty threshold that filters out surface-level memorization. The result is a benchmark that rewards genuine reasoning over recall.

The diagram below traces how a question earns its place in the benchmark.

Three-level taxonomy 10-option questions Difficulty threshold KINA 261 leaf disciplines A B C D E F G H I J chance ≈ 10% memorization filtered out genuine reasoning kept 261 disciplines high density living benchmark

Figure: how a KINA item is built — drawn from a three-level taxonomy of 261 disciplines, posed with 10 options to suppress guessing, and kept only if it survives the difficulty filter.

The annotation pipeline is designed around a Nash equilibrium incentive structure. Reviewer payoffs are structured so that honest, independent evaluation is the individually rational strategy, making collusion and lazy consensus irrational at equilibrium. This game-theoretic design is domain-agnostic and is released as a reusable framework alongside the dataset.

The two panels below contrast conventional blind-trust review with KINA’s incentive-aligned pipeline.

Blind-trust review (lazy consensus) KINA payoff design (equilibrium) question + draft answer R1 R2 R3 copies copies ✓ real check ✓ echoed ✓ echoed unanimous 3–0 label, only one independent evaluation question + draft answer R1 R2 R3 ✓ own verdict ✓ own verdict ✗ own verdict payoffs reward accurate, independent review — collusion doesn't pay honesty is the equilibrium strategy

Figure: why blind-trust annotation collapses into lazy consensus, and how KINA's reviewer payoffs make honest, independent evaluation each reviewer's rational strategy.

We evaluated 37 frontier models on KINA. Key findings include:

  • Significant room for improvement remains in domain-specific knowledge, even for top models.
  • Closed-source flagship models still maintain a leading position on knowledge-intensive tasks.
  • Web search tools exhibit a non-monotonic, U-shaped efficacy: both weaker and stronger models benefit substantially, while mid-tier models benefit less. The strongest model (Gemini-3.1-Pro-Preview) recorded the largest absolute gain at +5.17%.
  • The top-10 models show more pronounced performance differences in humanities and social sciences than in hard sciences.

The schematic below sketches the U-shaped search-tool effect from the third finding.

weaker models substantial gains mid-tier models smallest benefit strongest models largest gains +5.17% weaker stronger base model capability gain from web search illustrative, not to scale

Figure: schematic of the U-shaped efficacy of web-search tools — weak and strong models gain most, mid-tier models least; the largest absolute gain (+5.17%) went to Gemini-3.1-Pro-Preview.

KINA is designed to be continuously updated. The benchmark, annotation guidelines, and evaluation framework are released publicly alongside the dataset.