Browse
Claims
Search the graph by meaning. Each result carries its current verdict; open one to see its decomposition, provenance, and the reasoning behind the assessment.
ShowingImportancePrizesTopicAI Coding Benchmarks
Evaluation methods and benchmarks for assessing AI code generation and software development capabilities, including their validity, reliability, and correlation with real-world performance.
Coding benchmark scores and anecdotal reports overestimate real-world AI coding capability.
SupportedEvidence favors the claim, but the chain is incomplete or the sources are secondary.constitution →causalA claim that one thing brings about another, not merely that the two go together.constitution →Artificial intelligenceimportance · notableImportance 0.60, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution →
In GitHub's 2022 controlled experiment, developers using Copilot completed a coding task about 55% faster.
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitution →empirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitution →AI Developer ProductivityGitHub Copilot 2022 speed experimentGenerative AI productivity effectsimportance · minorImportance 0.30, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution →
AI coding benchmark scores are inflated by data contamination and memorization.
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitution →empirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitution →Benchmark data contaminationLarge language modelsimportance · notableImportance 0.50, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution →
Contribute
If a claim here is wrong, or missing evidence, open it: every claim page carries its own entry for challenges, evidence, and corrections. If the graph is missing a claim entirely, propose it below. A proposal is reviewed on its merits; accepted claims are matched against the graph and enter it with their reasoning on record.