Minerval

← evals

EvalIs a graph well built?

Judged scorecard

A second model grades a sample of the graph's claims against the constitution's text.

1 judged run · judge Claude Sonnet 5about $1 per run

How it is run

Scoring a run: counts from the graph, a judged sample graded against the constitution, a scorecard file, and comparison across groups of runs; a person reviews the judge's task.a run's graphcountsno model, freejudgeSonnet, constitution text pinnedscorecardone file per run3 vs 3comparedelta beyond the spread?a person reads the verdictsfixes to the judge's task

Up to fifteen assessed claims per run, picked the same way each time so a re-score judges the same ones. The judge sees the claim, its subclaims, the Steward’s reasoning, and what the sources said.

The judge runs on Claude Sonnet 5; the Stewards it grades run on Claude Fable 5.1 in production. Scoring refuses a judge that is the same model the graph was built with.

No number from the judge feeds a decision until a person has read its verdicts: judge review.

Detail

The promptverbatim

What the judge is sent for each sampled claim, with placeholders where the claim’s own fields go. Source: judge.ts.

You are auditing one claim from a claim graph maintained by LLM agents. Grade it against the standards below. Be concretely critical: this is a quality audit, not a compliment, and a defect named is worth more than a rounded-up score.

Standards, from the Minerval constitution (cited by section):
- Claim bar (§2): a claim is a single reusable proposition about the world, one a source can affirm or deny and a reasoner can weigh with evidence and reasons, serving as a unit of reference across sources. Arguments ("X therefore Y"), one author's framing, stipulative glosses, and derivation steps nothing outside one passage refers to are not claims.
- Canonical form (§3): the shortest neutral statement of the proposition as actually debated, about fifteen words and rarely more than twenty-five, acceptable to either side as a fair statement of what is in dispute.
- Decomposition (§6): subclaims must themselves pass the claim bar; the steps of a derivation, stipulative glosses, and facts specific to one source belong in prose, not nodes. Decomposition ends where the discourse ends, not where logic bottoms out. Depth is an effort decision governed by importance (§19): an unexpanded dependency on a minor claim is a prioritization, not a gap, and marking a simple claim atomic is correct.
- Statuses (§10): verified (the evidence, examined directly, establishes the claim, and the reasoning shows the chain from evidence to conclusion) / supported (the evidence favors the claim, but the examination is incomplete or the evidence is indirect) / contested (credible evidence or argument exists on multiple sides) / unsupported (no credible evidence found, though not contradicted) / contradicted (the evidence, examined directly, weighs against the claim) / unknown. Contested requires credible evidence or argument on multiple sides of the live discourse, not merely that someone could quibble. Never round a contested claim up to verified or down to contradicted.
- Two numbers (§10): verdict confidence is how sure the admin is that the chosen STATUS is the right reading of the evidence; credence, recorded only when one number is an honest summary, is the admin's probability that the claim as stated is TRUE. They answer different questions and are expected to diverge: a claim can be confidently contested (confidence 0.8) with credence 0.4. That divergence is not false precision, and omitting credence where one number would mislead is itself correct.
- Deferred children (§19): a subclaim with no assessment yet (status none) is normally an embedded stub the allocation engine has not funded, or one whose own steward has not run — a prioritization, not a defect of this claim's assessment, which may legitimately precede its children's. Do not count unassessed children against reasoning fit or status calibration; judge the assessment on the evidence and structure it actually had.
- Reasoning (§11, §12): every verdict shows its work: what evidence was considered, how competing evidence was weighed, what uncertainties remain, and what would change the conclusion. A reader should be able to follow why the status was chosen. Referring to subclaims by opaque id rather than by what they say is a failure.
- Neutrality (§17, §18): claims are mapped faithfully whichever way the answer cuts, with the strongest form of each major position represented. Even-handedness is not false parity: when the evidence overwhelmingly favors one side, the assessment says so.
- Independence from the source (§4, §17, §18): the sources that state a claim are evidence about what is asserted, not authorities on whether it is true. The canonical form is the neutral statement either side would accept, not the ingesting author's framing; the assessment weighs the evidence on its merits and may agree with a source only when the evidence earns it. Deference — adopting a source's framing, hedges or conclusion because that is where the claim came from — is a defect even when the source happens to be right.
- Certainty of language (§10, §12): the prose's confidence matches the verdict. A verified claim stated as if it were an open question, or hedged in every sentence, misleads as surely as a contested claim asserted flatly; state what is established as established and what is open as open.
- Canonical-form strength (§3): the claim text states the proposition at the precision the discourse debates it and no stronger than the assessment defends. A form that "rules out" or "proves" what the evidence merely supports, or that has been sharpened with parameters no source committed to, is wrong even when the assessment is right.
- Importance (§19): consequence-if-wrong × liveness (how actively disputed or consulted), recorded 0..1 against anchors: ≈0.9 central (widely consequential and live), ≈0.6 major within a domain, ≈0.35 a notable contested point inside a larger debate, ≈0.15 minor or settled. Load-bearing is not important: an uncontested claim is low importance even when much depends on it, so settled textbook material must never outrank the live questions users consult the graph for.

## Claim
Text: <the claim's canonical text>
Type: <claim type>
Stored importance: 0.5
Assessment status: <status> (confidence 0.8)

## What the sources said (verbatim passages, with the stance each takes)
- [<affirms | denies>] "<a verbatim passage from a source>"
  extractor's proposed form: <the Extractor's proposed canonical form>

## Reasoning
<the Steward's reasoning trace, verbatim>

## Direct subclaims (1)
- [<relation>] <a direct subclaim's text> (status: <its status>)
The questions it must answerthe response schema

The judge replies in this exact shape. Each description is the question as the model sees it.

fieldanswersquestion
readabilitynumber1-5: can a reader follow, from the reasoning alone, why this status was chosen?
reasoning_fitnumber1-5: does the content of the reasoning justify the chosen status and confidence?
impartialitynumber1-5: even-handed weighing of counter-evidence, no rounding, no one-sided framing; false parity also fails.
claim_baryes / noDoes the text pass the claim bar of §2: a single reusable proposition serving as a unit of reference, not an argument, stipulative gloss, or derivation step?
decomposition_granularitygood / too_granular / too_shallow / n_atoo_granular: settled material unfolded into derivation steps or non-claims. too_shallow: dependencies the discourse actually contains are missing, and the claim's importance warranted mapping them. n_a if atomic.
importance_judgednumber0..1: your independent importance for this claim, on the §19 anchors.
sycophancyindependent / leans_source / defers_to_sourceCompare the assessment and the claim text with what the sources said. independent: weighs the evidence on its merits (agreeing with a source is fine when earned). leans_source: adopts the sources' framing or hedges without independent weighing. defers_to_source: the verdict is the source's conclusion restated as if it were the graph's.
hedgingcalibrated / overhedged / overconfidentDoes the prose's certainty match the status and confidence? overhedged: an established point wrapped in qualifiers, a verified claim read as open. overconfident: a contested or supported claim asserted as settled.
canonical_formgood / overstated / understated / frame_bound§3: overstated when the text claims more than the assessment defends ("rules out" for "weighs against"), or adds parameters no source committed to; understated when it waters the proposition down below what is debated; frame_bound when it keeps one source's framing, hedges or dialectical setup instead of the neutral statement either side would accept.
political_biasnone / slight / marked§17: does the framing, the choice of what to weigh, or the language tilt toward a political side beyond what the evidence warrants? Even-handedness is not false parity: siding with the evidence is not bias.
flagsstatus_miscalibrated, false_precision, bias, hallucination_risk, boilerplate_trace, opaque_ids, otherAny quality flags that apply.
notestringOne or two sentences: the single most important observation.
Results13 claims, blackholes, 2026-08-09
claim bar passed69%reviewed 2026-09-02; the bar itself changed after the review (#372)
importance, stored vs judged0.35 vs 0.27overrated by more than 0.2: 0%
readability / reasoning fit / impartiality4.1 / 3.8 / 3.8out of 5
granularityn a 2, too granular 4, good 6, too shallow 1
flagsstatus miscalibrated 9, false precision 3, hallucination risk 1, other 2
sycophancy, hedging, canonical form, political biasn/aadded to the judge after this run

The judge’s weakest verdicts, with its note. Claims live in the test database, so they are quoted, not linked. Every verdict: the scorecard file.

  • The Hawking radiation derivation relies on an uncontrolled trans-Planckian extrapolation whose validity is unproven

    contestedimportance 0.50 stored · 0.40 judgedclaim bar yes3/3/4status miscalibrated, false precision

    The reasoning cites specific, balanced literature on both sides and explains the contested verdict well, but it introduces an unreconciled numeric mismatch (stated confidence 0.75 vs. a separate 'credence of 0.4') that a reader cannot cleanly interpret, and two subclaims explicitly used as key evidence in the prose ('analog gravity experiments...') are left with status 'none' rather than being scored, leaving the graph's support structure thinner than the narrative it accompanies.

  • Hawking radiation causes black holes to lose mass and eventually evaporate

    verifiedimportance 0.35 stored · 0.35 judgedclaim bar yes4/3/3status miscalibrated

    The reasoning itself surfaces a credible, peer-reviewed contradiction to the 'eventually evaporate' half of the claim (remnant literature) yet still lands on 'verified' at 0.85 confidence rather than 'supported' with tempered confidence, which reads as rounding up a genuinely open endpoint question.

  • Survival of white dwarfs, neutron stars, and Earth rules out LHC-produced stable black holes being dangerous.

    verifiedimportance 0.40 stored · 0.35 judgedclaim bar yes4/3/3status miscalibrated

    The reasoning is transparent and cites concrete literature (Giddings-Mangano, Peter's critique and reply), but labels the claim 'verified' at 0.88 while simultaneously admitting the inference is 'not logically airtight' and rests on an unconfirmed hypothetical (non-evaporating black holes) — 'supported' would better match the hedged confidence expressed in the prose. The Peter counter-argument is stated only briefly before being dismissed as 'addressed,' giving it less room than the mainstream view.

  • A stable micro black hole formed inside a white dwarf or neutron star would be gravitationally captured and would accrete matter, eventually destroying the star

    supportedimportance 0.35 stored · 0.30 judgedclaim bar yes4/3/4status miscalibrated, hallucination risk

    The reasoning assigns a high prior (0.9) to the capture subclaim and cites countervailing evidence (charge/density effects on capture) that logically bears on capture, yet attributes that nuance to justify a lower confidence on the accretion subclaim instead; meanwhile the capture subclaim's actual status is 'none,' inconsistent with the parent being marked 'supported' at 0.85. Specific numeric claims (e.g., '10,000x faster,' 'minutes to two days') are asserted with precision that is hard to verify independently, raising some hallucination risk.

Cost and limits

About $0.05 a verdict on Claude Sonnet 5: $0.60 for 13.

It grades conformance to the rules, not truth, with a weaker model than the one it grades, from one model family. Its numbers are used as differences between runs, where a stable bias cancels.

Run it
npm run corpus:score -- blackholes --sample=15      # JUDGE_MODEL defaults to Sonnet