Brainis runs on Brainis. Our missions are compiled here, our agents work inside our own products, and our deliverables pass through the same verification layer we sell. That means we have a rejection record, and publishing it is the most useful thing we can do with it.
One thing before the report. We are not publishing counts.
What you’ll learn
- Why the numbers are absent and when they arrive
- The five rejection categories that recur
- What the rejections say about where AI work actually breaks
- The two changes this month's shape produced
Why no numbers
The obvious way to write this post is with a table of counts and a trend line. We are not doing that, for the same reason the verified-outcomes counter on our homepage renders live records or nothing at all.
Every number that will ever appear there is a verification record, not a marketing estimate. Until that record is being rendered live, a count in a blog post would be a number sourced from a spreadsheet somebody maintained, presented with the authority of a system. That is the exact move this whole pillar exists to argue against, and doing it here would be the fastest possible way to be wrong about our own product.
So this is a report on shape. Categories, patterns, and what changed as a result. The counts follow when they can come from the ledger.
Important: Treat any vendor's dogfood metrics with the same suspicion you would treat their case studies, ours included. The question to ask is not what the number is. It is where the number comes from and whether it could have been edited.
The five categories
Criteria drift. The largest category, consistently. The deliverable is good work against a slightly different reading of the task than the acceptance criteria specified. This is the failure independent verification exists for, and it is the one self-review never catches — the producing agent assessed against its own reading and was satisfied.
Unsupported specificity. A number, a date, or a name appears in output with no traceable source. Frequently the figure is plausible and roughly right, which makes it more dangerous rather than less. Our criteria treat any unsourced specific as a rejection regardless of whether it happens to be correct.
Stale state. The work was correct against state that had moved by the time it was produced. This is a system problem more than a model problem, and its presence in the rejection record is a useful signal about freshness in a particular domain.
Scope creep inside a mission. The deliverable does what was asked and also three adjacent things nobody asked for. Rejected not because the extras are bad but because they were not scoped, budgeted, or checked, and shipping unscoped work quietly is how a governed system stops being governed.
Tone and claim violations. Copy that overstates. In our case that means our own claim registry: a sentence that asserts more than the registered claim permits gets rejected before a human ever reads it. Our public claims are checked in the same pass as everything else, which is the only way that discipline survives volume.
What the shape says
Three observations from reading a month of this.
The distribution is stable, and that is the point. The categories recur in roughly the same proportion week to week. A stable rejection profile is what makes a domain legible: you know what kind of thing goes wrong there, so you know what to strengthen. A domain whose profile is churning is a domain where something upstream is changing.
Zero rejections is a warning. We had one domain with a clean week and treated it as a problem rather than a success. Either the criteria are too loose to catch anything, or the volume is too low to mean anything. Both were true, as it turned out.
Criteria quality dominates model quality. Almost every improvement we made this month was to acceptance criteria, not to prompts or model selection. Your own acceptance criteria become verification checks, authored once and composed into libraries, and the library is where the leverage is. A sharper criterion improves every future run in that domain; a sharper prompt improves one.
The two changes
We tightened the unsourced-specific rule. It previously applied to numbers. It now applies to dates and named entities as well, after a run produced a confidently wrong date that was structurally identical to the number failures we were already catching.
We moved one domain down a rung. Its rejection rate was climbing over three weeks while its volume held steady. Under earned autonomy the record moves the level in both directions, so it dropped a level and stays there until the trend reverses. What earned autonomy means in practice covers why the downward move matters more than the upward one.
What we would tell you to copy
If you are running AI work at any volume, three things from this are portable and none of them require our product.
Write acceptance criteria before the work, specific enough that someone other than the author can check them. Categorise your rejections, because the categories are more informative than the rate. And treat a clean week as a question rather than an achievement.
The verification layer itself is on Verify. Our incident classes, published responses, and refund policy are on trust. The counts arrive here when the ledger can produce them, and not before.
Brainis Team
Notes on the company loop — company state, decisions, governed autonomy and verified work — from the people building Brainis and running on it.