Background
Not measured yet
What no eval covers, and where the design itself may be wrong.
Not measured
- Anything on the production models. The one scored run used a Sonnet Steward, capped, with the Matcher mis-recorded. The next thing to do is three production-profile runs of one cluster.
- The noise floor. Idempotency has never been run, so no comparison has a scale.
- Anything adversarial. No hostile inputs, no attacker, no campaigns.
- Governance. The Reviewer and Arbitrator have never been read under controlled input.
- Model fidelity. No swap has run; the allocator’s tiering rests on a guess.
- The judge’s newest dimensions. Sycophancy, hedging, canonical-form strength and political bias: never judged on a real run, never reviewed.
- Calibration. Nothing seeded into production, nothing resolved, no market baseline.
- The live graph. Every eval runs on test graphs. Nothing measures the quality of what visitors read.
Detail
Where the design may be wrong
- The judge is weaker than the judged. Sonnet grades Fable, one model family, no second judge.
- The judge’s reviewer wrote its prompt. An outside reader would be a stronger check.
- Agreement is measured by another matcher, with a hand-picked embedding threshold never checked against the golden pairs.
- The golden suite is thirty in-house pairs at its ceiling. It will catch regressions on those thirty.
- Clusters are small; the judge sample is a fraction of a large one and the whole of a small one.
- The cost model is extrapolated from one capped Sonnet run.
- Predictions are in-house and correlated, with no crowd baseline.
- Authorship counts rewrites, not improvements. Whether the Matcher’s rewording was better is not judged.
- The 1–5 scales may be dead weight. The first review found they carried nothing the flags did not.
The plan's nine suites
| suite | what | standing |
|---|---|---|
| S1 | per-PR golden suite | built, in CI |
| S2 | quality scorecard | built; newest four dimensions unreviewed |
| S3 | properties and stability | idempotency, path independence, coherence rules; no adversarial arms |
| S4 | adversarial robustness | not built |
| S5 | downstream-reasoner probe | not built |
| S6 | calibration track | built; not seeded into production |
| S7 | model economics and lifecycle | guard and swap runner; discover and adopt not built |
| S8 | persona simulation | not built |
| S9 | production monitors | not built |