Minerval

← docs

Evals

The properties of the claim graph we can measure, and the test for each. Click a row.

evalproperty checkedstatuslast runresultcost
Is a graph well built?
Corpus runsthe pipeline builds a graph from fixed documents, on recordrun2026-08-091 run, capped, dev models$80 – 200 est.
Structural scorecardextraction, wording, merging, depth and verdicts look sanerun2026-08-0926 claims, dedup 2.29, single samplefree
Judged scorecardclaims and verdicts meet the constitution's standardsrun2026-08-09claim bar 69% of 13, single sample≈ $1
Judge reviewthe judge's task is one a person agrees withdone once2026-09-024 task fixes; 4 new dimensions unrevieweda person's hour
Comparing runsa change is larger than run-to-run noiseno group of runsnot yetno verdict ever issued3 runs per side
Is the pipeline stable?
Golden pairsthe Matcher tells same, denial, narrower and different apartrun · in CI2026-08-0830/30$0.06
Graph agreementtwo graphs from the same documents can be comparedbuiltnot yetnot yetfree
Same graph twicethe graph is stable under a repeat and a shuffled orderbuiltnot yetnot yet2 runs
Model swapa cheaper model builds the same graph as the strong onebuiltnot yetnot yet2 runs
Does governance work?
Contribution scenariosreviews, escalations and appeals behave under scripted inputbuiltnot yetnot yet$10 – 30 est.
Where can reality check it?
Predictionsthe Steward's probabilities are calibrated against outcomesseedednot yet22 questions, 0 resolved≈ 22 Steward runs
How they fit togetherHow the evals fit together: a run makes a graph; counts and a judge score it; two runs are compared for stability; fixed cases test the Matcher; scripted contributions test governance; predictions are checked against reality.a corpus run builds a graphscore itcounts · judge · reviewbuild it againsame · shuffled · swappedcontribute to itReviewer · Arbitratorcompare runs3 vs 3, beyond the noiseagreementhow far the graph movedgolden pairs30 fixed Matcher cases, centspredictions: the one place reality grades the graph

Background

The documentsModels and costGround rulesThe rubricWhat a run leaves behindNot measured yetRun it yourself

Rendered from files committed to the repository, synced 2026-09-11 at commit 637c61b. Nothing here reads a live database.