FLF Epistack · submission notes
How Minerval handles the ingestion, structure, and assessment questions
Minerval is a public knowledge graph of claims. It reads sources, extracts the propositions they assert, links each claim to every source that speaks to it, decomposes a claim into what it rests on, and assesses how well the evidence supports it. The graph is maintained by LLM administrators bound by a public constitution, and every judgment carries a reasoning trace that anyone can inspect and challenge.
The best way to evaluate the system is to read the graph and the governing documents directly, so this page points at them first and then answers the brief’s questions in terms of what the system does today. Where we have not built something, or have built it but not yet tested it under pressure, the text says so plainly.
The three FLF case studies are the origin of SARS-CoV-2, the safety of micro black holes at the LHC, and the health effects of eggs. All three are now live in the graph, and the linked claims below are drawn from each.
The brief splits an epistemic investigation into three layers: ingestion, structure, and assessment. We find the same division useful, and the agent organization is built along it. What follows takes the brief’s questions in that order.
Layer 1 · Ingestion
Turning a messy, multi-source evidence base into something structured
Extract and attribute claims to specific sources, with provenance metadata (who said what, when, in what context).
The extractor reads a source and surfaces the discrete propositions it asserts. Each is recorded as an instance: the source’s own wording, the surrounding context, and whether the source affirms or denies the claim. The instances for a claim sit together on the claim page, so a reader sees every source that has spoken to the same proposition, in that source’s own words, in one place.
When a statement in a source is matched to a canonical claim, the admin creates an instance linking the utterance, with its original text and context, to the canonical claim. This preserves exactly what was said while enabling aggregation across sources.Administrator Constitution, §4
See it: the zoonosis-origin claim carries several sourced instances on its claim page.
What we do not do yet. We do not trace provenance recursively, following a claim back through the chain of who cited whom to its ultimate origin. Doing that exhaustively takes a great deal of recursive search, and it is not necessary for judging whether a claim is true, so we have not built it. It is clearly valuable and it belongs on the claim page, and we intend to pursue it. It should get easier as Minerval scales and more original sources are already in the graph to trace back to.
Identify when the same claim appears across multiple sources in different forms.
This is the matcher’s job, and the reason canonical forms exist. When a new source states a proposition the graph already holds, under different words or as its negation, the matcher links a new instance to the existing claim rather than minting a duplicate. The canonical form is kept short and frame-independent so that two authors arguing opposite sides of a question land on the same node.
A claim and its denial are not two claims but one. They pose the same question and turn on the same considerations, differing only in which answer a source endorses.Administrator Constitution, §2
See it: “SARS-CoV-2 has a laboratory origin” and the Huanan-market spillover claim are held as two live, opposed claims, neither merged into the other: lab origin, zoonosis.
Search for resources with bearing on the topics and subtopics at hand.
The Claim Steward, the agent that owns a claim, searches the web for evidence when it assesses. That is real, but it is assessment-time search aimed at getting one claim right, not a systematic survey of a field. Its limits are worth being honest about, and they are the subject of “surface what is missing” in the assessment layer below.
Capture useful metadata tags, relating sources and claims to topics, methodologies, deference, and assumptions.
We do not attach flat topic or methodology tags to claims. The graph is, in a sense, already metadata: the decomposition of a claim into its arguments, its subclaims, and its assumptions is recursively rich structure that a reader can follow as far down as they want. Rather than label a claim from the outside, we record the relationships that place it. An assumption, for instance, is not a tag but a subclaim on an assumes edge, which means it can itself be examined and assessed.
Claims decompose into other claims. The admin’s central structural function is to identify and articulate these relationships faithfully.Administrator Constitution, §6
Layer 2 · Structure
Documenting the relationships so the shape of the argument is navigable
Resolve the inference structure: which claims and evidence are offered as support for which other claims.
Each claim is decomposed into the subclaims it rests on, grouped under named arguments, with every edge labelled by how the child bears on the parent: requires, supports, contradicts, specifies, defines, or assumes. Each argument also carries a short written form stating how its subclaims combine. Decomposition stops at contestedness, not at logical bedrock, so a claim no informed person disputes is left a leaf even when much depends on it.
Decomposition ends where the discourse ends, not where logic bottoms out.Administrator Constitution, §6
See it: the furin-cleavage-site claim decomposes into two named arguments over six subclaims on its page and its map.
The written form states an inference without judging it. The judgment lives beside it, in the argument’s evaluation: the steward’s standing verdict on whether the inference goes through and which premises it lives or dies on. Because the evaluation tracks those premises as their own assessments change, a reader sees not just how an argument is arranged but how well it currently holds.
The judgment the written form withholds lives beside it, in the argument’s evaluation.Administrator Constitution, §7
See it: the LHC black-hole safety claim carries five named arguments, four for and one against, each separately evaluated, on its page and its map.
Represent the discourse structure: where people address different sub-questions, and their differences of emphasis.
When a claim has more than one distinct line of reasoning, each is a named argument with its own subclaims, for or against. Where the validity of an argument’s framework is itself in dispute, that meta-claim enters the same structure as an assumption, so a dispute about the terms of the debate lives in the claim layer rather than off to the side. Different sources emphasize different arguments, and because each source’s instances record which claims it engaged and on which side, the emphasis is recoverable.
A claim may have several distinct arguments: coherent, self-contained lines of reasoning that bear on its truth. Each argument groups its own subclaims; different arguments may share subclaims while arranging them differently, or rest on different premises entirely.Administrator Constitution, §7
Capture relationships regarding similar but not identical claims: different framings, conditions, caveats, or estimates of uncertainty.
Two statements are the same claim only when they turn on the same considerations, so a canonical form is pinned to what is actually in dispute. “Inflation was high” meaning “above two percent” is a different claim from the same words meaning “above wage growth,” and the two are held apart and related rather than merged. Claims that differ only by a condition or a caveat are linked, often by a specifies edge, so the qualification is visible as structure.
Two superficially identical statements may be different claims if they turn on different considerations; two differently phrased statements may be the same claim if they differ only in wording.Administrator Constitution, §3
See it: the egg cluster keeps the general claim that regular egg consumption raises cardiovascular risk apart from the narrower finding that the association appears at high but not moderate intake, the caveat held as its own node rather than folded in.
Track how the structure evolves over time.
Assessments are provisional and versioned. Each claim keeps its assessment history, and when a subclaim’s assessment changes, the steward of a dependent claim is notified to reconsider. When the pipeline’s own rules change in a way that would alter what gets minted or how it is valued, a claim’s pipeline epoch lets an older cohort be retired and re-derived rather than silently carried forward.
The world changes: new evidence emerges, studies are retracted, predictions come due. The admin updates assessments when the underlying situation changes.Administrator Constitution, §22
What we do not do yet. The history we keep is mainly at the assessment level. We do not yet present a full structural diff over time, a view of how the shape of an argument itself changed as claims were added, merged, or split.
Layer 3 · Assessment
Evaluating what to believe, and what to look at next
Identify rhetorical moves that carry more persuasive weight than evidential weight.
We handle this structurally rather than by trying to detect rhetoric as such. Three things do the work. Canonicalization strips a source’s framing down to a neutral proposition that both sides would accept, so a persuasive framing does not survive into the claim. The division of roles means the agent that reads the source and might be swayed by it, the extractor, is not the agent that judges whether the claim is true. And the Claim Steward assesses the canonical claim on its evidence, outside the original source’s framing, under a constitution that binds it to weigh evidence rather than authority or presentation.
The admin assesses claims on the merits. Where a source is relevant, the admin opens it and reads it whole: the methods, the data, the reasoning, not the abstract and the headline.Administrator Constitution, §9
Flag correlated evidence being treated as independent.
We do not have a mechanism aimed specifically at this, and the constitution does not call it out. In the graph so far we have not seen the mistake, and we think the emphasis on mapping the logical arguments and weighing the evidence fairly tends to guard against it. We are not against telling a model what to watch for, but we have found that naively adding rules for how to think can cause overcorrection, so we are reluctant to add a warning here before we have evidence that it helps.
Identify cruxes: the disagreements that, if resolved, would most change the overall picture.
This is close to the center of the design. A claim’s importance is defined as roughly how consequential it would be to get wrong, multiplied by how genuinely contested it is, judged against all of claimspace rather than the local neighborhood. That is a crux measure directly: a claim that much turns on and that is actually in dispute. The steward spends effort in proportion, and decomposition drives toward the specific subclaims where the disagreement actually lives.
What earns high importance is that getting the claim wrong would be consequential and the claim is contested or heavily consulted: a live crux, not settled scaffolding.Administrator Constitution, §19
See it: “SARS-CoV-2 has a laboratory origin” is scored as a central, contested crux on its page.
Surface what is missing: important sources or perspectives not represented in the working knowledge base.
Honestly, what is most often missing is more sources. For SARS-CoV-2 we ingested only a corner of the discourse, centered on the Rootclaim debate and mainstream scientific writing. That is a reasonable corner, since it holds itself to higher standards and tends to focus on the questions that matter, but we cannot rule out that it left out claims which, properly investigated, would add arguments to the central question or change how the graph weighs it.
The steward does search proactively, but that only goes so far. An agent trying to work out what is true is naturally drawn to the arguments it knows best from pretraining and trusts most, which in a case like this one tends to sanewash, and which leaves out claims that are in fact routinely made. The agents are told to map the arguments in good faith and keep an open mind, but they still carry a bias toward consensus and share many of the blind spots of the existing epistemic environment. Casting a wider net at ingestion is part of the answer. The other part is that Minerval is built to be open to the public, so that the perspectives an agent misses can be contributed and then held to the same standard as everything else.
The admin makes the structure of the disagreement visible, so that users can see what a claim rests on, where consensus exists and where it does not, and whether each point of disagreement is empirical, and so potentially resolvable with evidence, or reflects differences of values or definitions.Administrator Constitution, §1
Provide frameworks for calibrating confidence that account for out-of-model error, adversarial information environments, and the limits of any single analyst.
The status vocabulary is about the state of the evidence rather than a bare probability. A claim is verified, supported, contested, unsupported, contradicted, or unknown, each defined by how the evidence stands. Two numbers may sit beside a status, answering different questions: a verdict confidence, always recorded, is how sure the steward is that the status is the right reading of the evidence; a credence, recorded only where a single number is an honest summary, is the probability that the claim itself is true. A claim can be confidently contested, the disagreement near-certainly real while nobody knows the answer, and where one number would be false precision the credence is omitted, which is itself information. Effort scales with importance, and before it records a verdict the steward runs a second pass that tries to refute its own conclusion. On the limits of any single analyst, our answer is the organization itself: separate agents with separate duties, coherence sweeps that check assessments against each other across the graph, and heavier scrutiny reserved for the most important contested claims.
The admin does not round uncertain claims up to “verified” or down to “false.” The graph’s value comes from honest representation of the state of knowledge.Administrator Constitution, §10
See it: the egg-consumption claim is held confidently contested, with a verdict confidence of 0.85 while its credence sits at 0.33, on its page.
What we do not claim. We have not built explicit modeling of out-of-model error, and we would not claim to have solved calibration. The multi-agent design is our current answer to the single-analyst limit, and we expect to keep revising it.
Distinguish what the debate settled from what it merely performed settling.
Refusing false resolution is the system’s first commitment. A genuinely contested question is held contested, with the strongest form of each side represented, rather than rounded to a winner. At the same time, not every disagreement is genuine, so a fringe or bad-faith objection is noted without being raised to false parity. The SARS-CoV-2 origin question is a fair test: the graph holds the laboratory-origin claim and the market- spillover claim as two live contested claims, neither merged into the other and neither quietly resolved, with the weight of expert opinion recorded without erasing the minority case.
An admin who clearly maps an unresolvable disagreement has done their job well. An admin who imposes false resolution has failed, and so has an admin who withholds a well-supported verdict out of misplaced even-handedness.Administrator Constitution, §1
See it: the LHC black-hole safety claim is assessed supported, while the objection that survives scrutiny is recorded as a methodological dispute about proof under small probabilities of catastrophe, not as evidence of danger.
What we have not done, and where we want to learn
The honest boundary
Some of the machinery the brief asks about is deployed but not yet proven. Contributions, conflict review, escalation, arbitration, and an audit function that samples decisions and watches for coordinated manipulation are all built and running, but we have not tested them against real bad-faith contributions at scale. We describe that layer as something we expect to work and are still hardening, not something we have demonstrated. The graph does not yet field dedicated adversarial agents that probe public contributions for manipulation, and we think it will need them to be robust; the steward’s self-refuting second pass and the audit role are a start, not a finish. On this subproblem we do not claim to be ahead of people who have worked on it more directly, and we are looking to learn from them.
One design choice is worth stating, because it is deliberate rather than unfinished. We did not build assessment as a judge scoring two advocates in a staged debate. We think a single Claim Steward with a duty to the truth and to good epistemics is likelier to reach a sound verdict than a judge refereeing adversaries. The lawyering model, for all that it beats the known alternatives in a courtroom, rewards rhetorical sleight of hand, which is the opposite of what we want at the point of judgment. Adversarial agents have a place in stress-testing a verdict, not in standing between the evidence and the verdict.
Provenance tracing, a structural diff of a claim over time, and topic-level metadata are things whose value we see and have chosen not to build yet. The reasoning behind each is in its section above.
Temporary page for the FLF Epistack competition. It describes the system as it runs today and links live claims; the graph is maintained continuously, so a linked assessment may have moved since this was written. Read the graph at /claims, the governing texts under /docs, and the code on GitHub.