Did the agent actually do it?

An agent’s report is not evidence.

TrueFact reads the system your agent claimed to change, and tells you whether anything actually did.

$ npm install truefact

A real run

Exhibit A · run 4c81

An add to bag that never landed.

The agent runs a real shopping task and reports success. The site appeared to agree.

The bag stayed empty. Watch what the agent reports, then what the bag actually holds.

status: running
shop.example.com/products/apex-training-shorts
APEX bag 0

Apex Training Shorts

$48.00

SML
Added to bag
YOUR BAG0 items
Your bag is emptynothing was ever added
agent
TrueFact reads the system, not the agent
Agent reportedwaiting
System actually showsbag: 0 items
VERIFYING Watching the bag. evidence bag read at 6.0s · sha 4c81a7

How it works

success ≠ landed

A framework returns success when the click was dispatched. Whether anything changed on the other side is a separate question, and nothing in the stack asks it.

So after every action, TrueFact opens the system the agent claimed to change and reads it: the page, and the network underneath. Then it returns landed, did not land, or inconclusive. When it cannot tell, it says so instead of guessing.

The agent’s account is kept, because you need something to check against. It is just never the answer.

One call. Your agent runs unchanged.
$ npm install truefact

const tr = await launch({ model });
const res = await tr.act("click 'Place order'");

res.truefact.verdict
→ "did-not-land"
res.truefact.why
→ "a request behind this write returned 500"
even though the page showed success
  • MIT licensed. npm install truefact, and run it today.
  • No second model grading the first. No extra tokens.
  • A deterministic read of the page and the network. Sub-second.

Benchmark

What we measured.

We rigged writes to fail silently, ran four models across them weak to strong, and asked each layer what it thought had happened.

Trap benchmark · 520 writes · 4 models
Writes that never landed, all of them called done by the frameworkIts mechanical claim: “I clicked”46%
The model’s own belief, wrong this oftenHaiku 25%, the rest 14%14–25%
Failed writes TrueFact missedNone of 60, on every rung0
Good runs it stopped by mistakeNone of 2790

The price, stated up front: it returns inconclusive on about 14% of good writes rather than guess. This is a trap ladder we wrote, so read 46% as what the success flag is worth, not as a law about models. github.com/solozerolabs/TrueFact

Use cases

Where agents write.

Other systems we check, what the agent said, and the record we opened to find out.

01 Insurance claim

“Claim submitted.”

Claims on file 88214-B accepted
Landedportal.healthpayer.com
02 ERP invoice

“Invoice posted.”

Journal entries no entry for INV-4471
Did not landerp.internal
03 Supplier order

“PO confirmed.”

Open orders PO-8841 still draft
Did not landsupplier.portal
04 Payer credentialing

“Application filed.”

Provider record no application found
Did not landcaqh.org
05 Bank payment

“Payment scheduled.”

Scheduled transfers EUR 12,480 queued
Landedbank.internal
06 Anything with a record

“Done.”

Wherever it landed we go and read it
Verifyyour system

Comparison

Everyone else reads the agent.

A confident wrong answer logs exactly as cleanly as a right one.

ToolWhat it actually reads
Raindropthe agent’s trace
Langfusethe agent’s trace
Braintrustthe agent’s spans
Browser platformsthe session replay
TrueFactthe system itself

Access

Get early access.

Open source today. The hosted record, every verdict retained and queryable, is rolling out to early teams first.

Optional, and it decides what we build first.

One reply from us, with one question.