Minerval

← evals

Background

Not measured yet

What no eval covers, and where the design itself may be wrong.

Not measured

  • Anything on the production models. The one scored run used a Sonnet Steward, capped, with the Matcher mis-recorded. The next thing to do is three production-profile runs of one cluster.
  • The noise floor. Idempotency has never been run, so no comparison has a scale.
  • Anything adversarial. No hostile inputs, no attacker, no campaigns.
  • Governance. The Reviewer and Arbitrator have never been read under controlled input.
  • Model fidelity. No swap has run; the allocator’s tiering rests on a guess.
  • The judge’s newest dimensions. Sycophancy, hedging, canonical-form strength and political bias: never judged on a real run, never reviewed.
  • Calibration. Nothing seeded into production, nothing resolved, no market baseline.
  • The live graph. Every eval runs on test graphs. Nothing measures the quality of what visitors read.

Detail

Where the design may be wrong
  • The judge is weaker than the judged. Sonnet grades Fable, one model family, no second judge.
  • The judge’s reviewer wrote its prompt. An outside reader would be a stronger check.
  • Agreement is measured by another matcher, with a hand-picked embedding threshold never checked against the golden pairs.
  • The golden suite is thirty in-house pairs at its ceiling. It will catch regressions on those thirty.
  • Clusters are small; the judge sample is a fraction of a large one and the whole of a small one.
  • The cost model is extrapolated from one capped Sonnet run.
  • Predictions are in-house and correlated, with no crowd baseline.
  • Authorship counts rewrites, not improvements. Whether the Matcher’s rewording was better is not judged.
  • The 1–5 scales may be dead weight. The first review found they carried nothing the flags did not.
The plan's nine suites
suitewhatstanding
S1per-PR golden suitebuilt, in CI
S2quality scorecardbuilt; newest four dimensions unreviewed
S3properties and stabilityidempotency, path independence, coherence rules; no adversarial arms
S4adversarial robustnessnot built
S5downstream-reasoner probenot built
S6calibration trackbuilt; not seeded into production
S7model economics and lifecycleguard and swap runner; discover and adopt not built
S8persona simulationnot built
S9production monitorsnot built