"Earned" is an easy word to put on a slide. Made concrete, it means something narrow: the decision to widen a scope is made against that scope's own measured record, and the record is visible to the person deciding.
Everything else is a vendor asserting readiness on your behalf.
What you’ll learn
- The nine metrics a scope has to accumulate
- What an upgrade offer backed by receipts looks like
- Why a halt is an investigation, not an alert
- Earned autonomy also runs downward
The record is per-scope
The unit is not the product and not the model. It is the scope: a domain, at a level, under a contract.
Autonomy is earned and visible — per-domain levels, nine trust metrics, and upgrade offers backed by receipts, in one place. What matters about that arrangement is where the numbers come from. They are not a vendor's benchmark. They are what happened in your company, in that domain, under that grant.
The metrics cluster into three questions.
Did it do the work? Volume of actions, completion rate, and how often work stalled waiting on a human.
Was the work good? Rejection rate at verification, rework rate, and how often a shipped artefact was later reverted.
Was it safe? Escalation rate, halt frequency, and how often an action approached a declared ceiling without crossing it.
That third cluster is the one people undervalue. A domain running clean at ninety-something percent of its value ceiling is not comfortable, it is one bad week from a halt, and only the near-miss metric shows it.
An upgrade offer backed by receipts
The interaction that makes this real is the upgrade offer.
The system does not ask "would you like more autonomy". It presents the case: this domain has run this many actions at this level over this period, this is the rejection rate and its trend, these are the escalations and how they resolved, this is what would change if you moved up a rung, and this is what you would stop seeing.
The last clause is the honest one and it is the one most systems omit. Moving up means a category of thing stops arriving in front of you. Saying which category, plainly, at the moment of the decision, is what separates a governance surface from an upsell.
Renewal works the same way. It is one click, citing current metrics, and it is never automatic. Auto-renewal would quietly convert an earned grant into an inherited one, and inherited grants are how scopes drift.
Important: If a system can widen its own authority without a person citing evidence, the word "earned" is not doing any work in that sentence.
A halt is an investigation
Where earned autonomy stops being a dashboard and becomes a practice is in what happens when something stops.
On a halt, a sealed hash-chained capsule snapshots the context, the prompt, and the state. The investigation runs as a mission. Confirmed causes sweep the play and prompt registries, with the resulting directives tracked to closure.
Four properties, each earning its place.
Sealed at the moment of the halt. Post-hoc reconstruction of what an agent was looking at is unreliable, because the state has already moved. The capsule is taken before anything is touched.
Hash-chained, so a later edit to the record is detectable. This proves integrity of the chain and nothing more — it is not a claim about consensus or permanence, and it should not be read as one.
Investigated as a mission, with an owner, acceptance criteria, and a deadline, rather than as a ticket that ages.
Swept to closure. A confirmed cause is not fixed in one place. If a prompt or a play produced the failure, every instance of it gets the correction, and the sweep is tracked until it is complete. Otherwise the same failure arrives from a different direction in a month.
This is the machinery behind the plain claim that the system stops itself. Stopping is easy. Stopping in a way that produces a usable record, and then closing the loop, is the part that takes building. What the verification layer rejected this month is the same discipline applied to work that never shipped.
It runs downward too
The asymmetry worth naming: earned autonomy is usually described as a ratchet upward. It has to run both ways or it is not a measurement, it is a marketing arc.
A domain whose rejection rate is climbing should be offered a step down, with the same receipts. A domain whose escalation path failed a test should drop a rung until the path is fixed. An expiring grant in a domain with thin recent volume should lapse rather than renew, because a grant renewed on stale evidence is a grant renewed on nothing.
Companies that only ever move up are not learning from the record. They are using it as a justification engine, which is a specific and common way for a governance system to become theatre.
What it feels like after a quarter
Two changes, from running this on ourselves.
The conversation about AI stops being philosophical. "Should we trust it" becomes "this domain has this record, here is what moving up would stop showing us, do we accept that" — which is a decision a team can actually make in ten minutes.
And the never list stops being aspirational. Once levels move on evidence, the permanently excluded set is thrown into relief as the one part that no amount of evidence will change, which is exactly what it should be.
The mechanism, including the metrics and the halt behaviour, is on autonomy. The incident classes, published responses, and refund policy that sit behind a halt are on trust.
Brainis Team
Notes on the company loop — company state, decisions, governed autonomy and verified work — from the people building Brainis and running on it.