Minerval

← evals

Background

The documents

Four fixed sets of documents, each chosen to stress something different.

The four clusters

clusterdocumentswordschosen for
blackholes43,380a settled question with a deep argument: the control case
eggs31,617a mundane question where the dispute is about methods, not facts
lableak53,319the most contested case, both sides steelmanned
lethalities1188,567dense overlap between documents that argue with each other

Detail

blackholes4 documents

The LHC micro black hole safety case — one of the three FLF Epistack case studies (flf.org/epistack-competition). A mostly-closed, near-uncontested question ('Will CERN's Large Hadron Collider create a black hole that destroys the Earth?') that nevertheless rests on a large body of interacting physics. Chosen to stress how the system handles a settled-but-deep argument: heavy claim overlap between institutional reassurance, the rigorous safety derivation, and a published dissent, plus a few genuinely contested cruxes (does Hawking radiation exist? are the astrophysical bounds airtight?). Sources are ordered so the institutional and foundational safety arguments are ingested before the dissent that contests them.

Curated, committed markdown assembled from public sources: CERN's LHC safety pages and the LHC Safety Assessment Group (LSAG) report; Wikipedia's 'Safety of high-energy particle collision experiments' (CC BY-SA 4.0); and two peer-reviewed/arXiv papers (Giddings & Mangano 2008; Plaga 2008). Each post's `url` records its provenance. This is a `web` cluster: the markdown is the pinned source of truth and is not refetched.

eggs3 documents

The health effects of eggs — one of the three FLF Epistack case studies (flf.org/epistack-competition). A deliberately mundane, open-ended question ('Are eggs good or bad to eat? For whom? How would we know?') that is really a question about ways of knowing: dietary cholesterol vs blood cholesterol, observational cohorts vs randomized trials, confounding and reverse causation, whole-diet context, and subgroup effects (notably diabetics). Chosen to stress how the system handles genuinely unresolved, methodology-driven disagreement where the cruxes are about evidence quality rather than a single fact. Sources span the mainstream-nuanced view, the conflicting meta-analytic evidence, and the guideline/epistemics angle.

Curated, committed markdown assembled from public sources: the Harvard T.H. Chan School of Public Health 'Nutrition Source' page on eggs; Wikipedia's 'Egg as food' health/nutrition sections (CC BY-SA 4.0); and a synthesis of the dietary-cholesterol guideline history and study-design debate. Each post's `url` records its provenance. This is a `web` cluster: the markdown is the pinned source of truth and is not refetched.

lableak5 documents

The origin of SARS-CoV-2 (zoonosis vs. lab leak) — the hardest and most-contested of the three FLF Epistack case studies (flf.org/epistack-competition). A live, high-stakes dispute with a rich public record, anchored by the 2024 $100k Rootclaim debate (Saar Wilf vs. Peter Miller) that two expert judges ruled decisively for zoonosis, and where six independent Bayesian analyses of the same evidence spanned 23 orders of magnitude. Chosen to stress contested-claim handling, head-to-head disagreement, and cruxes at maximum difficulty. Sources are ordered overview → zoonosis case → zoonosis primary source → lab-leak case → lab-leak primary source, and both sides are steelmanned so the graph must hold genuine disagreement rather than collapse to one side. Each side's synthesis is followed by a faithful representation of a real primary source (Worobey et al. 2022 for zoonosis; the 2018 DEFUSE proposal for lab leak) so the load-bearing cruxes carry independent, citable instances rather than resting on the debate write-up alone.

Curated, committed markdown assembled from public sources: Scott Alexander's writeup of the Rootclaim COVID-origins debate (the FLF starting material) and the FLF case description for the framing; Wikipedia's coverage of the origin investigations (CC BY-SA 4.0) and the published market-origin analyses (Worobey et al. and Pekar et al., Science 2022) for the zoonosis case; and the documented lab-leak arguments (WIV proximity, the DEFUSE proposal, the furin cleavage site, intelligence assessments) for the lab-leak case. Each post's `url` records its provenance. This is a `web` cluster: the markdown is the pinned source of truth and is not refetched.

lethalities11 documents

The 2022 'List of Lethalities' AI-risk debate. Yudkowsky's AGI Ruin anchor post plus direct responses, the sharp-left-turn sub-thread, and the basic-x-risk-case sub-thread. Chosen for dense claim overlap and head-to-head disagreement, to stress disambiguation, canonicalization, and contested-claim handling. Posts are ordered so the anchor and foundational posts are ingested before the responses that argue against them.

lesswrong.com GraphQL API (contents.markdown); IDs are pinned for reproducibility

  • AGI Ruin: A List of Lethalities · Eliezer Yudkowsky · anchor
  • Where I agree and disagree with Eliezer · Paul Christiano · response-to-anchor
  • My Objections to "We’re All Gonna Die with Eliezer Yudkowsky" · Quintin Pope · response-to-anchor
  • DeepMind alignment team opinions on AGI ruin arguments · Victoria Krakovna et al. · response-to-anchor
  • A central AI alignment problem: capabilities generalization, and the sharp left turn · Nate Soares · sharp-left-turn-foundational
  • Refining the Sharp Left Turn threat model, part 1: claims and mechanisms · Victoria Krakovna et al. · response-to-sharp-left-turn
  • “Sharp Left Turn” discourse: An opinionated review · Steven Byrnes · review-of-sharp-left-turn
  • Counterarguments to the basic AI x-risk case · Katja Grace · x-risk-case-foundational
  • Response to Katja Grace's AI x-risk counterarguments · Erik Jenner · response-to-x-risk-case
  • Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover · Ajeya Cotra · takeover-model
  • Six Dimensions of Operational Adequacy in AGI Projects · Eliezer Yudkowsky · adequacy-framework
Isolation

Runs use a separate database. Every harness tool refuses the live database by name. So the claims a run produces are not the claims on this site, and cannot be linked from here.