Verified AI work

Evidence graphs: proving AI work happened

A deliverable is a claim. The evidence graph is what backs it: which agent produced it, under whose authority, against which criteria, checked by what, at which revision.

5 min readVerified AI work

Once software does work, a new question appears that did not exist when people did all of it: how do you know it happened the way the system says it happened.

With people, the answer is social. You ask them. With agents operating at volume, that does not scale and would not be reliable if it did. The replacement is an evidence graph — a structured record of what was produced, by what, under what authority, checked against what, and at which revision.

What you’ll learn

  • The six nodes every piece of governed work leaves behind
  • Why the seal has to pin a revision
  • What independent re-verification actually proves, and what it does not
  • Turning the graph into something an auditor can read

Six nodes

Each completed piece of work leaves a small graph. The nodes are unremarkable individually and useful together.

The intent. The objective this work serves, traced back to the strategy that produced the mission. Without this node, the work cannot answer "why".

The authority. The grant under which the action was permitted: scope, level, who signed it, when it expires. Every consequential action is governed by an authority contract and is auditable, and the graph is where "auditable" becomes concrete.

The producer. Which agent, which model, which version, which prompt lineage, which tools it was permitted to call. Access is scoped — each agent gets only the context and tools its mission allows — so the producer node also records what it could not reach.

The artefact, as a diff. Not just the output but the change: old versus new, with inline threads and findings on exact lines. Every deliverable arrives as a reviewable diff before it counts as done.

The criteria. The acceptance criteria in force at the time, in the version in force at the time. Criteria change, and a check is only interpretable against the criteria that were actually applied.

The verdict. The verification result, from an independent checker, pinned to a specific revision.

The seal pins the revision

This one property does more work than the rest combined.

The verified seal renders only for the exact revision that was checked. Changed work shows "verified at rev N, changed since" rather than continuing to display a pass.

The reason is that a seal which survives edits is actively harmful. It transfers confidence from a checked artefact to an unchecked one, silently, in the direction that matters. Anyone reading it is more wrong than they would have been with no seal at all, because they now have a reason to skip looking.

"Verified at rev N, changed since" is an uncomfortable state to render. It says the system does not currently vouch for what you are looking at. That discomfort is the correct output, and the honesty of the whole layer rests on being willing to display it.

What re-verification proves

Records can be exported, and an exported decision record can be independently re-verified against the chain. Paste the record and the verifier re-walks its hash chain in front of you, link by link.

It is worth being exact about what that demonstrates, because this is an area where claims routinely inflate.

What it proves: integrity. Each link's recorded hash is recomputed and compared. If a field in the record was altered after the fact, the walk breaks at the altered link and the break is visible. Tampering is detectable.

What it does not prove: that the record was ever published anywhere, that it was witnessed by a third party, or that it cannot be deleted. Integrity of a chain is a narrower property than permanence, and conflating them would be exactly the kind of overstatement this pillar exists to reject.

Today the publicly checkable set is three labelled sample records, one of which is deliberately broken so that detection can be demonstrated rather than asserted. You can try the walk yourself on verify a record. A non-public record is indistinguishable from a non-existent one, by design — there is no way to enumerate other companies' records through the verifier.

Tip: When any system claims tamper-evidence, ask to see it fail. A verifier that has never shown you a broken record has not shown you anything.

From graph to compliance pack

The practical payoff arrives when someone external asks a question.

Because the graph exists as a by-product of doing the work rather than as a reporting exercise, the answer to "show me how this decision was made and who authorised it" is a query rather than a project. Controls mapped, evidence attached, exported as a pack — generated by the same systems that ran the company, not assembled afterwards by someone reconstructing intent from calendar invites.

Two cautions on that. It is a pack of evidence, not a certification, and nothing about producing one implies an external attestation you do not hold. And the pack is only as good as the criteria that were in force when the work ran — a graph is a faithful record of a weak check as readily as a strong one.

The verification layer and how criteria compose into libraries are on Verify. Why the checker has to be independent of the producer is in the agent that does the work shouldn't grade it.

An artefact says what it is. The graph says how it came to be, and only one of those survives being questioned.

Verified AI workVerify

Brainis Team

Notes on the company loop — company state, decisions, governed autonomy and verified work — from the people building Brainis and running on it.

Bring Brainis the company you want to build.

Start from an idea. Connect what exists. Either way, leave with the next move — and a system that delivers it.