Evals
The properties of the claim graph we can measure, and the test for each. Click a row.
| eval | property checked | status | last run | result | cost |
|---|---|---|---|---|---|
| Is a graph well built? | |||||
| Corpus runs | the pipeline builds a graph from fixed documents, on record | run | 2026-08-09 | 1 run, capped, dev models | $80 – 200 est. |
| Structural scorecard | extraction, wording, merging, depth and verdicts look sane | run | 2026-08-09 | 26 claims, dedup 2.29, single sample | free |
| Judged scorecard | claims and verdicts meet the constitution's standards | run | 2026-08-09 | claim bar 69% of 13, single sample | ≈ $1 |
| Judge review | the judge's task is one a person agrees with | done once | 2026-09-02 | 4 task fixes; 4 new dimensions unreviewed | a person's hour |
| Comparing runs | a change is larger than run-to-run noise | no group of runs | not yet | no verdict ever issued | 3 runs per side |
| Is the pipeline stable? | |||||
| Golden pairs | the Matcher tells same, denial, narrower and different apart | run · in CI | 2026-08-08 | 30/30 | $0.06 |
| Graph agreement | two graphs from the same documents can be compared | built | not yet | not yet | free |
| Same graph twice | the graph is stable under a repeat and a shuffled order | built | not yet | not yet | 2 runs |
| Model swap | a cheaper model builds the same graph as the strong one | built | not yet | not yet | 2 runs |
| Does governance work? | |||||
| Contribution scenarios | reviews, escalations and appeals behave under scripted input | built | not yet | not yet | $10 – 30 est. |
| Where can reality check it? | |||||
| Predictions | the Steward's probabilities are calibrated against outcomes | seeded | not yet | 22 questions, 0 resolved | ≈ 22 Steward runs |
How they fit together
Background
The documentsModels and costGround rulesThe rubricWhat a run leaves behindNot measured yetRun it yourself
Rendered from files committed to the repository, synced 2026-09-11 at commit 637c61b. Nothing here reads a live database.