Independent AI output verification
We check what your AI says. Independently.
Your AI system makes claims all day. Which ones survive being attacked? The Verification Lab runs your AI outputs through adversarial checks by multiple independent model families, none of which built your system, and sends back a written evidence report: what held, what broke, and the full receipt trail.
Not useful? Tell us where it fell short and we send half your money back. A complete report is published in full, findings and receipt trail, so you can judge the work before you buy. Our method's benchmark scores are public too.
Two ways in
Checking after the fact, or clearance before you publish
Both run on the same method and both hand back the same thing: a dated, per-item record you can show someone else. They differ in when they sit in your process, and that changes what you get.
Verification Lab · after
You already have AI outputs. We attack them with multiple independent model families, none of which built your system, and return a written evidence report: what held, what broke, and the receipt trail.
Choose this when the reader is a conformity file, a procurement review, an auditor or an insurer.
Facts Or Fibs? · before
You have copy about to ship. Every checkable claim comes back passed, held, or flagged for a person, each with a dated receipt written to sit in a substantiation file.
Choose this when the reader is the ASA, a client, or your own compliance sign-off, and the question is whether a claim can be defended before it runs.
The problem, and the dates
Your conformity file needs evidence, not assurances
Most EU AI Act rules apply from 2 August 2026. High-risk-system rules generally apply later: 2 December 2027 for Annex III uses and 2 August 2028 for systems embedded in regulated products. Your conformity assessment needs documented evidence that the system was tested against the applicable requirements.
Your AI vendor can test its own system. We add an independent challenge to that evidence.
Where we fit, said up front rather than in the small print: we produce the technical testing evidence your internal compliance team or your notified body needs. We are not the notified body, we do not sign off your conformity, and you would not want a supplier who claimed otherwise. You keep the declaration and the CE marking. We give you defensible evidence to put behind them.
Why independence is the product
A grader with something to defend is not an independent check
When the company that built your system also grades it, the grader has something riding on the result. We hold no model to defend and no vendor to please.
- Every claim is attacked by multiple independent model families, never the one that produced it.
- Checkers work independently: they do not see other checkers' verdicts before returning their own.
- Every verdict ships with its receipt trail: what was asked, what was attacked, what survived, what did not. You can hand the whole trail to your auditor.
- We tested our cross-family checking panel on a public benchmark and publish every item it missed. Client work stays confidential unless the client says otherwise.
- We build our tools all the way down. The methods behind these reports run on a verification language we develop in-house, and every claim we make about how it behaves gets proved or refuted by machine rather than taken on trust.
What "independent model family" actually means here
It is a specific test, not a marketing phrase. Two checkers count as independent families in our method only if all of the following hold:
- Different vendor. Separate companies, separate commercial interests, separate legal entities.
- Separately trained weights. Not two sizes of the same model, not a fine-tune of the other, not the same base with a different wrapper.
- No shared inference path. Independent endpoints, so one provider outage or one bad deployment cannot silently take out both opinions.
- Blind to each other. No checker sees another's verdict or reasoning before returning its own.
- Never the system under test. If the model that produced your output belongs to one of the families, that family is removed from the panel for your job.
We do not publish which specific vendors sit in the panel, for the same reason a testing house does not publish its instrument serial numbers: it invites gaming and it changes as models change. What we do publish is the rule above, the measured error rate of the method, and every verdict with its reasoning, so you can judge the checking rather than the branding. If your own compliance process requires the vendor list, ask us and we will tell you in writing under NDA.
Our own evidence · measured, published, gated
We hold our method to a number, and we publish the attempt that failed
Anyone can say their checking is rigorous. Here is ours, measured against a bar we set before scoring and cannot move afterwards: how often we wrongly flag a TRUE claim as false, which is the error that wastes your time. We fixed the ceiling at 10%.
False flags, certified run
0
out of 98 true claims decided
Statistical upper bound
4.1%
98.3% one-sided, under our 10% ceiling
Falsehoods caught
99%
counting every hand-off as a miss
Sent to a human
1.5%
where the checkers disagreed
"Your certificate is on your data, not mine. So what am I actually buying?"
Fair, and it is the right question. Two honest answers.
First, the certified class is the one your outputs mostly live in. What we certified is checking whether a factual claim is true, on claims validated against named sources. That is what an assistant answering questions, a summariser, or a model-written report mostly produces. The corpus where the same method does not certify is a set of adversarial trick questions written by researchers to fool models into confident wrong answers. That is a harder and different job, and we would rather show you the number than hide the corpus.
Second, nobody can certify on your data before seeing it, so we made it cheap to find out. That is exactly what the £95 Spot Check is for: ten of your real outputs, three working days, and if the report is not useful you tell us where it fell short and we send half back. You are not being asked to trust a number from someone else's corpus. You are being asked to spend £95 to get the number on yours, and we carry half the downside if it disappoints you.
What that certificate covers, and what it does not
It covers fresh, source-validated, stable-fact claims across eight subject areas. It does not cover adversarial trick-question benchmarks, where the same method measures 8.5% and is not certified, and it does not cover your documents until we have run them. Certification of a checking method is always certification on a stated class of claims. Any vendor telling you otherwise is selling the word, not the measurement.
So why hire us rather than anyone else? Not because we are flawless, and the numbers above are the proof we are not. Because when your auditor, your insurer or your customer asks how reliable the checking was, you can hand them a figure with a method and a stated class behind it. Ask any other supplier of AI assurance for their false-positive rate on a named corpus and see what comes back. We publish ours, we publish the corpus where the same method does not clear the bar, and we publish the experiment that failed. That is the whole product: not a promise that the machine is right, but evidence about how often it is wrong, in a form you can file.
Our first attempt at this certification failed, and we published that too: the way we selected test questions accidentally filtered them down to ones every model already answered correctly, so the result was technically valid and practically worthless. We retired it, fixed the selection, and ran it again. The full benchmark, both evals, and every item our method missed are public.
How it works
Four steps, and a person signs the report
01 · Send the outputs
A transcript set, a batch of answers, a model-written report, claims your system makes in production. Tell us what the system is supposed to do.
02 · We attack them
A fixed adversarial sequence: factual grounding, internal consistency, refutation attempts, edge probing. Independent families, blind checks.
03 · You get the evidence report
Written findings in fix-first order, plus the machine receipt trail as an annex, formatted to drop into a conformity or audit file.
04 · You fix, we re-check
Re-verification of remediated items is included in the larger tiers.
Exactly what to send, so nothing stalls after you pay
Reply to your Stripe receipt, or mail hello@thefoundworks.com, with these four things. There is no form to fill in and nothing to install.
- The outputs. Paste them, or attach a file. Number them if you can, so our findings line up with your list.
- What the system is supposed to do. One or two sentences. "It answers customer questions about our returns policy" is enough.
- Who it talks to. Customers, staff, regulators, the public. This changes what counts as a serious error.
- Anything you already suspect. Optional, and we check the rest regardless.
We reply to confirm within one working day and the turnaround clock starts from that confirmation, not from your payment. If you have not heard from us in one working day, chase us and we will treat the delay as ours.
Pricing
Fixed prices, published turnarounds
No quotes, no calls, no retainer. Secure checkout by Stripe. After you pay, send your outputs and brief to hello@thefoundworks.com, or reply to your Stripe receipt, and the clock starts.
What counts as one output or claim: a single discrete thing your system produced or asserted. One answer, one paragraph, one summary, one recommendation. A long document is not one output, it is however many separate claims we have to check inside it. If you are not sure how your material divides up, send it before you buy and we will tell you which tier it needs. We would rather answer that email than take the wrong payment.
Spot Check · 3 working days
£95
Up to 10 outputs or claims. The single biggest reliability risk we find, the checks each claim survived and failed, one concrete fix to make first.
Standard Verification · 5 working days · most chosen
£245
Up to 50 outputs or claims. Full adversarial battery, written findings in fix-first order, receipt-trail annex, one re-check of fixed items within two weeks.
Evidence Pack · 10 working days
£495
Everything in Standard plus a documentation-ready annex mapped to your conformity file's testing-evidence section, a methods statement your auditor can read, and a follow-up re-verification round.
Facts Or Fibs? · clearance before you publish
Substantiation has to happen before you publish. Checking has not kept up.
The CAP Code requires an advertiser to hold evidence for a claim before it runs. "Clinically proven", "97% of customers", "the UK's fastest", "made with recycled materials". Each one has to be defensible if the ASA asks.
Two things changed. Volume: generative tools collapsed the cost of producing copy, so far more of it ships, across more channels, faster than manual review can follow. Provenance: when a draft is machine-written, plausible statistics and attributions appear that nobody chose to make, and the person approving the copy often cannot say where a number came from.
So the old control, the copywriter knows their sources, has quietly stopped holding, and nothing has replaced it. What is left is either hourly legal review, applied to a fraction of output, or a general-purpose AI checker whose own reliability nobody has measured. You cannot put "we asked a chatbot" in a substantiation file.
How the checking differs from the Lab's: the method is the same, the unit is not. The Lab reports on a body of outputs and tells you what broke under attack. Facts Or Fibs decides each claim separately, and every claim is judged blind by checkers from different independent model families. Where they agree, that is the verdict. Where they disagree, it goes to a person, because disagreement is information rather than something to average away.
01 · Send the copy
Campaign copy, a product page, an email sequence, a script. Before it ships.
02 · We find the claims
What is actually a factual assertion, as opposed to puffery, which is not a claim and is not our business.
03 · Each claim is cleared
Passed, held, or flagged for a person, per claim, not per document.
04 · You keep the receipt
What was asked, who checked it, what each returned, what settled it. Written to sit in a substantiation file.
What our published numbers are, and what they are not
The scores above measure the method, on fresh source-validated factual claims. They are not a measurement of Facts Or Fibs on advertising copy, because that product is in development, has no customers, and we have not measured it. Advertising claims are compressed, implied and often carried by images, and our honest expectation is that the number there will be worse. When we have it, we will publish it, including if it disappoints us.
We have form on this: our first attempt at certifying the method passed and we retired it anyway, because its acceptance gate had accidentally selected only questions every model already answered.
Read this before you rely on the clearance path
We are not a regulator and we do not approve advertising. Nothing we produce is legal advice or a compliance sign-off. The advertiser remains responsible for substantiating its claims, and always will be. If anyone offers to take that responsibility off you, be suspicious of them.
We also do not check everything. Claims carried by images, by implication, or by the overall impression of an advert are outside what we measure today, and we will say so on any report rather than let you assume otherwise.
The first pilots
We are looking for a small number of agencies and in-house marketing teams to run the first paid pilots: we clear up to 100 factual claims from one campaign and return them within one working day, each passed, held or flagged, with receipts.
You get claims cleared cheaply and a look at whether the receipts are any use to you. We get the only thing that actually matters at this stage: whether this is worth paying for on real work. If the report is not useful, tell us where it fell short.
What we are, and what we are not
Read this before you buy
We are an independent verification lab. We produce technical evidence about how your AI system's outputs behave under adversarial checking.
We are not a notified body. We do not carry out third-party conformity assessment, issue certificates of conformity or EU type-examination certificates, or sign off compliance, and nothing we produce is legal advice.
To be precise about where we sit, because the distinction matters: under EU product law you affix your own CE marking and draw up your own declaration of conformity. A notified body issues certificates where third-party assessment is required. We are neither. Our report is technical test evidence you can put in the documentation behind your declaration, whichever assessment route you use.
Your data: we process what you send solely to produce your report, we dont train anything on it, and we delete working copies within 30 days after delivery, or sooner on request.
Straight answers
The questions people actually ask
Cant my AI vendor do this? They can evaluate, and they should. What they cannot be is independent of their own system. The value of this report is exactly that nobody involved in the checking has anything riding on the answer.
What if our system passes everything? Then you have a dated, receipted record that it did, which is precisely what a testing-evidence section is for. A blind adversarial pass is designed to find issues worth fixing, if they are there.
Is this only for EU AI Act work? No. The dates above are one reason people come looking, but the report is the same evidence any procurement, insurance, or internal-audit process asks for.
Who reads the outputs? Automated adversarial checking with human review before anything ships to you. A person signs every report.
Start with a Spot Check
£95, up to 10 outputs or claims, back in 3 working days. If the report is not useful, tell us where it fell short. This is the feedback that improves the checks, so we refund you half for your trouble. Reply to your Stripe receipt within 14 days of delivery, tell us what missed, and we send £47.50 back. No form, no argument, and we keep the note.