Minerval

Browse

Claims

Search the graph by meaning. Each result carries its current verdict; open one to see its decomposition, provenance, and the reasoning behind the assessment.

ShowingImportancePrizesTopicUnit of randomization2 claims

Claims about which units of an experiment (individuals, tasks, clusters) were randomized and how that choice sets the effective sample size, statistical independence, and design effect for estimating a treatment effect — e.g., tasks-vs-participants randomization. Excludes generic critique of a study's findings or external validity.

METR's measured AI slowdown remained statistically significant when accounting for developer-level clustering.
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionMETR developer productivity experimentAI Developer Productivityimportance · minorImportance 0.30, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
Issue-level randomization made METR's 246 completed issues, not its 16 developers, the effective sample for estimating the AI slowdown.
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionMETR developer productivity experimentAI Developer Productivityimportance · minorImportance 0.35, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution

Contribute

If a claim here is wrong, or missing evidence, open it: every claim page carries its own entry for challenges, evidence, and corrections. If the graph is missing a claim entirely, propose it below. A proposal is reviewed on its merits; accepted claims are matched against the graph and enter it with their reasoning on record.