“Claim submitted.”
Did the agent actually do it?
TrueFact reads the system your agent claimed to change, and tells you whether anything actually did.
$ npm install truefact
A real run
The agent runs a real shopping task and reports success. The site appeared to agree.
The bag stayed empty. Watch what the agent reports, then what the bag actually holds.
$48.00
bag read at 6.0s · sha 4c81a7
How it works
A framework returns success when the click was dispatched. Whether anything changed on the other side is a separate question, and nothing in the stack asks it.
So after every action, TrueFact opens the system the agent claimed to change and reads it: the page, and the network underneath. Then it returns landed, did not land, or inconclusive. When it cannot tell, it says so instead of guessing.
The agent’s account is kept, because you need something to check against. It is just never the answer.
$ npm install truefact
const tr = await launch({ model });
const res = await tr.act("click 'Place order'");
res.truefact.verdict
→ "did-not-land"
res.truefact.why
→ "a request behind this write returned 500"
even though the page showed success
Benchmark
We rigged writes to fail silently, ran four models across them weak to strong, and asked each layer what it thought had happened.
The price, stated up front: it returns inconclusive on about 14% of good writes rather than guess. This is a trap ladder we wrote, so read 46% as what the success flag is worth, not as a law about models. github.com/solozerolabs/TrueFact
Use cases
Other systems we check, what the agent said, and the record we opened to find out.
Comparison
A confident wrong answer logs exactly as cleanly as a right one.
| Tool | What it actually reads |
|---|---|
| Raindrop | the agent’s trace |
| Langfuse | the agent’s trace |
| Braintrust | the agent’s spans |
| Browser platforms | the session replay |
| TrueFact | the system itself |
Access
Open source today. The hosted record, every verdict retained and queryable, is rolling out to early teams first.
One reply from us, with one question.