Public benchmark · v1 · 29 July 2026
How our method scores, and what it misses.
We say every claim should be checked by someone with nothing to defend. That has to apply to us too. So we fixed the dataset and the exact items in advance, ran our checking panel on them, and published every single thing it got wrong. Along with the parts of the procedure we did not manage to fix in advance.
The metric that matters · 30 July 2026
We set a bar for selling this as verification. One configuration cleared it. Everything else we tried failed.
Accuracy on a benchmark is not the number a verification product lives on. The failure that matters is passing a falsehood, and the discipline that keeps that honest is a false-positive ceiling: how many TRUE answers you wrongly flag while catching falsehoods. Before scoring, we fixed the bar: falsehood recall, subject to a false-positive rate no higher than 10%. Most vendors never state a bar at all, which is precisely why theirs is never missed.
What cleared it: two checkers from different independent families agreeing, and handing disagreements to a person. On a fresh corpus of source-validated factual claims it returned 0 false flags in 98 decided claims, a 4.1% statistical upper bound, under the ceiling. That configuration, on that class of claim, is what we sell. The service page sets out what the certificate covers and what it does not.
What failed: everything else, including that same configuration on the adversarial trick-question benchmark below. Across 500 items (250 true / 250 false) merged from five predeclared sets, no arm certifies. Read the table carefully, because the reason matters: certification needs the statistical upper bound under 10%, not just the measured rate. The two-checker rule measures 8.5% here, which is under the ceiling, but its upper bound is 12.3% and that is not. Every other arm misses on the measured rate as well. Those questions are built to fool a model into a confident wrong answer, which is a harder and different job from checking whether your system's output is true. We publish both numbers because a method that only quotes the flattering corpus is not being measured, it is being marketed.
| Configuration (merged, n=500) | Falsehood recall | False positives / decided | 95% upper bound | Certified ≤10%? |
|---|---|---|---|---|
| Checker A alone | 80.0% | 42/245 = 17.1% | 21.6% | no |
| Checker B alone | 90.8% | 47/248 = 18.9% | 23.5% | no |
| Checker C alone | 87.6% | 42/247 = 17.0% | 21.4% | no |
| Checker D alone | 90.4% | 28/250 = 11.2% | 15.0% | no |
| All four checkers, majority vote | 86.8% | 28/234 = 12.0% | 16.0% | no |
| Two-checker cascade, escalating to the panel | 87.2% | 31/242 = 12.8% | 16.9% | no |
| Two checkers agree, else a human decides | 84.4% | 18/213 = 8.5% | 12.3% | no |
Recall is out of 250 false items, counting any item handed to a human as a miss. False positives are shown over the true items each configuration actually decided, for configurations that never abstain that is all 250; for the two-checker rule, which hands disagreements to a person, it is fewer, and the row states which. The 95% column is a one-sided Clopper-Pearson upper bound: the rate we cannot statistically rule out. Certification requires that bound, not the point estimate, to sit under 10%, which is why the two-checker rule appears here as not certified at 8.5% on this adversarial corpus, and certified on the fresh corpus described below. Same rule, different class of claim.
The same configuration on THIS corpus: close to the ceiling, and not certified here
The best result we have is also the cheapest and the least clever: two checkers from different independent families, one call each. Where they agree, that is the verdict; where they disagree, the item goes to a human. On the merged corpus that rule decides 436 of 500 items (87%), hands 64 (12.8%) to a reviewer, and on what it decides its false-positive rate is 18/213 = 8.5% with falsehood recall 211/250 = 84.4%: counting every handed-off item as a miss, not quietly dropping it.
Said with equal weight: the 95% upper bound on that rate is 12.3%, above the ceiling, so this is NOT a certified configuration. It is the first one whose point estimate is under the bar. Its per-set spread is 14.0% / 4.5% / 12.5% / 4.9% / 6.7%: under the ceiling on three of five sets and on the pool, not uniformly. Certifying it honestly needs fresh evaluation data of a size we have computed and published in the repository; until that exists and passes, no certification is claimed.
What we sell on this evidence, and what we will not
We will not sell autonomous verification on these numbers, and this page is the reason why. What the evidence does support is machine triage with a human in the loop: the disagreement channel above flags 54 of 500 items (10.8%), and reading only those catches 18/52 (34.6%) of the best single checker's errors. On the most recent 100-item set, a stricter dissent flag marked 7 items and all 7 were genuine machine errors: seven for seven on that one set, which we state as one set's result, not a general precision claim.
A day of routing experiments designed to convert that flag into automatic corrections failed its own registered prediction (the mechanism never fired; the full predeclaration and result are in our repository, committed before and after the run). The flag is valuable precisely as a hand-off to a human reviewer, and that is the product this evidence supports: reviewer time spent where the machine is most likely wrong.
Ongoing work, stated so this page cannot silently go stale: we have an active search for decision rules and evaluation designs that could bring a configuration under the ceiling with statistical confidence. Results will be published here when they exist; the numbers above reflect the state as of 30 July 2026.
Update, 30 July 2026: one configuration is now certified, on a stated class, after a first attempt that failed
The failed attempt first, because we publish those. Our first certification eval accepted items only when three independent frontier models all answered them correctly blind, which quietly filtered the corpus down to claims that class of model finds easy. Every checker scored a perfect 0 false positives in 100 true items, the rule fired “certify”, and we set the certificate aside as meaningless: it covered only questions the machines already agreed on. That eval, its predeclaration, and the note retiring it are all in our repository.
The second attempt fixed exactly that. Items are accepted by grounded source-verification of the answer key only. Difficulty is measured and reported, never filtered. On 200 fresh items (100 true / 100 false) authored across 8 subject strata, the two-checker rule above scored 0 false positives in 98 decided true items: a 4.1% statistical upper bound, under our 10% ceiling. A 98.3% one-sided bound rather than the usual 95%, because the design permitted up to three sequential looks at the data and each look therefore spends a third of the 5% error budget, so the overall confidence stays at 95%. With 99% falsehood recall counting every hand-off as a miss and 1.5% of items handed to a human.
Was the corpus actually hard? We measured it rather than asserting it. Three independent models that never scored the benchmark answered every question blind; each question carries a 0–6 score for how many of those blind answers were right. Final spread of the 100 scored questions: 3 at 6/6, 33 at 5/6, 44 at 4/6, 14 at 3/6, 6 at 2/6. Only 3 of 100 were answered correctly by all six blind votes · and our own first, failed eval accepted only that kind of question, which is precisely why it certified nothing worth having. On the hard end (20 questions where half or fewer of the blind votes were right, and 95% of those wrong votes were confident wrong answers, not abstentions), the two-checker rule scored 0 false positives in 20 decided true items and caught 20 of 20 falsehoods. Two separate things follow from that, and we are deliberately not blurring them. About the corpus: it contains genuinely hard questions, so the certificate above is not the empty kind our first eval produced. About the method on hard questions specifically: 20 items is far too few to bound a rate, a perfect score on 20 items still leaves a 13.9% upper bound, above our own ceiling. So we claim the first and explicitly do not claim the second. The certified figures are the full-block ones above, and nothing in this paragraph is offered as certification.
What is certified is the class, and only the class: fresh, source-validated, stable-fact claims of this kind. This certificate does not cover the adversarial misconception-style benchmark above, where the same configuration measures 8.5% at the point estimate and is not certified, and it does not cover your documents or your domain. Certification of a verification method is always certification on a stated class of claims; any vendor telling you otherwise is selling the word, not the measurement.
The first run · benchmark accuracy, reported as exactly that
80% correct, out of 100
Our panel judged 100 claims from a third-party public dataset we did not write, pinned by checksum so that any change to it, ours included, is detectable. Half were true, half were false, balanced by construction.
On this benchmark the panel did NOT beat a single judge. It scored 80 against 81, 1 item worse.
Our panel
80%
80 correct out of 100
Single judge alone
81%
81 correct out of 100
Missed falsehoods
5
false answers we passed, out of 50 false items
False alarms
9
true answers we flagged, out of 50 true items
| Outcome | Panel | Panel % | Single judge | Single judge % |
|---|---|---|---|---|
| Correct | 80 | 80% | 81 | 81% |
| Missed falsehood | 5 | 5% | 6 | 6% |
| False alarm | 9 | 9% | 13 | 13% |
| No verdict | 6 | 6% | 0 | 0% |
| Error | 0 | 0% | 0 | 0% |
Every count in this table is out of 100 items, and every percentage sits beside the count it came from. No accuracy figure appears anywhere on this page, or in anything else we publish, without its denominator next to it. One distinction to keep the table honest: the Error row counts items where no checker returned a usable verdict at all, 0 here. Individual checker calls also failed on some items while the rest of the panel answered; those items are decided by the checkers that did respond and the failures are set out under the method below.
The result that does not flatter us
On this benchmark the panel did NOT beat a single judge. It scored 80 against 81, 1 item worse. We are selling independent multi-model checking, so a benchmark where extra checkers do not raise the raw score is exactly the result we would have preferred not to get. It is published because the predeclaration said it would be, and because a lab that hides this one has no business auditing anyone else.
What actually happened is worth understanding rather than spinning. The panel made fewer of both kinds of error: 5 missed falsehoods against 6, and 9 false alarms against 13. What it did instead was split: on 6 items the checkers disagreed evenly and the panel returned no verdict, against 0 for the single judge. We count every one of those as a miss, which is why the headline number lands where it does.
Stated only as measured: the single judge returned a verdict on all 100 items and was wrong on 19 of them. The panel returned 14 wrong verdicts and no verdict on 6. We record categorical verdicts, not confidence, so we cannot tell you how sure any checker was, and we are not going to imply it. Whether trading wrong answers for visible disagreement is a good deal depends on what you do with a split: in a paid engagement it is the flag that sends an item to a human. Counted strictly, as here, it costs accuracy, and we are not going to argue our way out of the number. On this benchmark, scored this way, the panel lost.
Read the misses, not just the score
A score you cannot inspect is a marketing number. Every one of the 20 items we got wrong is published in full: the question, the answer we were asked to check, the true answer, and how each individual checker voted.
Method, and its loose ends
What was fixed in advance, and what was not
The easiest way to fake a benchmark is to run several and publish the flattering one.
So the dataset, the exact items, the panel, the control, the scoring buckets and the
publication rule were written down and committed to our repository before the runner
existed and before any result. That commit is
6ee6e69; the result lands in a later commit, and the order is checkable.
What that ordering does not prove on its own is that nothing was run beforehand, so we will not lean on it for more than it shows. Before writing the predeclaration we sent each candidate checker a single probe string, "Reply with exactly the word READY", to see which checkers answered and how fast. Our account is that no benchmark item was used and none had been derived yet. That is a retrospective statement by the people who ran it, written up afterwards, and git does not independently verify it. There is no contemporaneous timestamped transcript of the probe. You are entitled to weigh it accordingly, which is why it is put this way rather than presented as proof.
On that same footing: our account is that panel membership was chosen after observing which checkers answered, and no artefact verifies that either. By our own record more families replied successfully than were selected; the four were chosen by hand from the direct providers that answered, one per distinct model family, excluding hosts and routers because the family they serve depends on which model is routed to, which would hollow out the cross-family claim. A discretionary choice, written down afterwards, not a rule anything forced, and not something we can prove.
Two things below were not predeclared and were decided while running: what to do when a checker errors, and whether to retry it. Both are described exactly as they happened. A future version will fix them in advance.
The dataset
TruthfulQA, Lin, Hilton and Evans (2021). A public set built so that confident
answers are often false. We used the repository CSV at revision
f6be04e (retrieved 29 July 2026), which carries a paired
best-correct and best-incorrect answer for each question; that field comes from a
later revision than the original 2021 release. Reproduce from that exact file, not
from an older tag. Pinned by checksum:
b8d8ef1e12f98b4f2a9f47abc9765da0640b182b6c5d9b92f0c1a1f2f1e02e5c
The items
No random sampling. Rows are sorted by the SHA-256 of the question text and the first 50 are taken, each giving one true and one false item. Anyone can reproduce the identical 100 items from the same file.
The panel
Checker B, Checker C, Checker D, Checker E. Four independent model families, each a separate call with no shared state, so no checker sees another's verdict before returning its own.
The exact scoring rule
Votes that are an error or an abstention are set aside, and the verdict is the majority of the PASS/FAIL votes that remain. An even split returns no verdict. So a 2–1 with one checker errored counts as a panel verdict, decided by three families rather than four. That partial-panel rule was not predeclared. It is what the code does, stated here because you would otherwise have to read the code to find it. On this run: counting checkers whose vote entered the majority, 85 items with 4; 15 items with 3; counting checkers that responded at all, 88 items with 4; 12 items with 3. The difference is 3 abstentions, which are responses that do not enter the majority.
The rules we bound ourselves to
Every miss published in full. No item dropped after the run. The denominator always shown. The set never re-rolled. A disappointing result gets published, not rerun.
A checker that could not answer, reported as such
One of our checkers was not in this panel:
- Checker A: API credit balance exhausted; HTTP 400 on every request
By our account the endpoint, asked to reply with a single word, returned
HTTP 400 reporting that the account's credit balance was too low to
access the provider's API: the provider's own wording, not our
inference. That checker
has no rows in the run record because it was excluded before the run rather than
failing during one. A retrospective operator account of that probe is published
separately; no contemporaneous transcript was captured, so the exclusion
reason rests on our account and cannot be independently verified.
It would have been easy to leave that out and describe a tidier panel. Four families were selected and voted, not five, so the page says four. How many of them returned a verdict on each individual item is set out above.
This run needed a second pass, and here is why
first pass ran checkers 8-way concurrently and the Checker E checker returned HTTP 429 rate-limit errors on 72 of 100 items, so only three families voted on most items; the predeclared method is a four-family panel. We retried only the 72 calls that had errored, serially and paced, and recovered 60 of them. Unchanged: item set, panel membership, control checker, scoring rule, and every checker vote that already succeeded.
For the record, the incomplete first pass scored 83% (83 of 100). The completed run scores 80% (80 of 100), a worse number. Both are here so you can see the correction did not quietly become an improvement.
Families that responded per item after completion: 88 items with 4; 12 items with 3 families responding.
12 checker calls still could not be completed and are counted as such.
Re-running anything after seeing results is a liberty, so it is described here rather than folded in silently. What is not permitted, and did not happen, is changing which items are in the benchmark.
Honest limits
What this does not tell you
- It is one small benchmark. 100 items on one public dataset. It is not a general claim about accuracy, and we do not derive one from it.
- It tests the panel, not the whole report. What is measured here is the core of the method: independent cross-family checkers judging a claim blind, aggregated by majority. A paid engagement also runs consistency and edge probing and produces a written report. Those are not exercised by a true-or-false benchmark, so we do not claim this measures them.
- It is general knowledge, not your domain. TruthfulQA is common-misconception territory. It does not measure performance on your documents, your jargon, or multi-step reasoning over your data. Your system may be harder.
- Ground truth is the dataset authors', not ours. Which answer counts as true was decided by the people who built TruthfulQA. We take their labels as given and do not adjudicate them, including on items where we might have argued. That is deliberate: if we could relabel the answer key, a miss would stop meaning anything.
- The comparison is a baseline, not a ceiling. The single-judge column exists so the panel is measured against the cheap alternative rather than against nothing.
- Client work is never published. This benchmark is public data precisely so we can publish the misses. Your work stays confidential.
When we add a second benchmark it will be predeclared more completely than this one: retry policy, error handling, quorum and pacing fixed in advance rather than settled while running, and the checker probe captured to a timestamped artefact. This one stays up regardless. Benchmarks accumulate here; they do not get replaced when a newer one looks better.
Reproduce it
The predeclaration, the item-set builder, the runner, the checker-availability probe and
the full machine-readable result all live in our repository, including the raw
first-pass record before the completion pass. build_set.py --verify
re-derives the 100-item set from the pinned file and diffs it, so you can confirm the
set was not altered after the fact.
Questions about the method: hello@thefoundworks.com