Minerval

Browse

Claims

Search the graph by meaning. Each result carries its current verdict; open one to see its decomposition, provenance, and the reasoning behind the assessment.

ShowingImportancePrizesTopicBenchmark data contamination1 claim

Leakage of benchmark or test-set items into model training data, allowing memorization rather than genuine capability to inflate evaluation scores; covers detection, measurement, mitigation, and evaluations on clean or held-out data across AI task domains.

AI coding benchmark scores are inflated by data contamination and memorization.
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionAI Coding BenchmarksLarge language modelsimportance · notableImportance 0.50, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution

Contribute

If a claim here is wrong, or missing evidence, open it: every claim page carries its own entry for challenges, evidence, and corrections. If the graph is missing a claim entirely, propose it below. A proposal is reviewed on its merits; accepted claims are matched against the graph and enter it with their reasoning on record.