Every company believes it learns from its decisions. Very few can show the evidence, because the evidence would have to be a record of what they expected, made before they knew, and almost nobody keeps one.
Without that record, review is an exercise in narration. The outcome is known, so the story assembles itself around the outcome, and everyone leaves with a lesson that feels earned and is mostly invented.
What you’ll learn
- What hindsight actually does to an organisation's memory
- The four fields a prediction record needs
- Why automatic grading matters more than good intentions
- What the record is worth once there is enough of it
Hindsight is not a bias, it is a rewrite
The familiar framing is that hindsight makes past events seem more predictable than they were. True, and understated. What actually happens is that the memory of the prior belief is replaced. People do not remember being uncertain and then discount it — they remember having expected roughly what occurred.
This is why post-mortems tend to converge on a tidy causal chain, and why the same organisation makes the same class of error repeatedly while sincerely believing it has learned from each instance. There was never a preserved prior to compare against, so there was never anything to be surprised by.
The correction is boring and effective: write the prediction down, in advance, in a place that cannot be edited afterwards.
Important: A prediction you can revise after the outcome lands is not a record. It is a draft of a story about yourself, and it will be revised in your favour without any dishonest intent whatsoever.
The four fields
A prediction record does not need to be elaborate. It needs four things.
The range. Not a point. The band you actually believe, with the relevant tail named.
The basis. The two or three assumptions the range rests on, stated well enough that someone could later check which one broke.
The decision it informed. Predictions made in the abstract do not teach you much. Predictions attached to the choice they justified teach you a great deal, because the interesting failure is usually the decision, not the forecast.
The version. Who or what produced it, at what version. A human estimator and a model at a given version are different predictors with different track records, and merging them makes both records useless.
Predictions carry ranges and a model version, and are preserved unchanged when reality arrives. That last clause is doing the load-bearing work. Preservation is the whole mechanism.
Grading has to be automatic
Here is the part every good-intentions version of this fails on.
Recording predictions is easy and slightly enjoyable. Going back months later to check them is neither, and it competes against everything else in a given week. So it does not happen, the file grows, and the file stops being read.
Every prediction is registered when it is made, and sealed. That half is running today and it is the half that defeats hindsight, because the record cannot be edited into agreement with the outcome.
Grading it automatically when the outcome lands is the other half, and we have not built it yet. Checking a prediction is still a human step here, which means this post is describing a discipline we hold ourselves to rather than a mechanism you can lean on. That is not a convenience feature we are deferring: it is the difference between a practice that survives contact with a busy quarter and one that does not, and it is why the automatic half is the next thing on this part of the roadmap.
Automatic grading would also remove the selection problem. When a person chooses which predictions to revisit, they revisit the memorable ones, and memorable correlates with surprising, and surprising skews the sample. A system that grades everything grades the boring ones too, and the boring ones are where calibration actually lives.
Predicted and actual outcomes land in the Decision Ledger and improve the next recommendation — the mechanism is described on memory.
What the record is worth
Three things, in increasing order of value.
Better arguments now. When a prior prediction exists, a disagreement can be settled by looking rather than by seniority. This benefit arrives immediately and is worth the effort on its own.
Known error shape. After enough graded predictions, patterns appear that no individual would have noticed. A team may be well calibrated on delivery timing and badly calibrated on adoption. Knowing which is which changes how much padding each kind of estimate deserves, and it is not knowable without the record.
A model that improves. Registered predictions, once graded, are the training signal for the prediction layer itself. This is the compounding part, and it is why the Business World Model treats the record as infrastructure rather than reporting. A model trained on text about business has read every case study and has never had one of its own calls scored.
Starting without any software
The practice does not require a system, and starting it manually is a good way to find out whether you will keep it.
Pick the decisions that clear a threshold — a spend, a hire, a commitment, a strategic bet. For each one, before deciding, write four lines: the range, the basis, the decision, the date you expect to know. Put them in one place, in a file nobody edits.
Two things will happen within a quarter. You will find at least one decision where writing the range down changed the decision, because stating the low end out loud made the downside real. And you will find at least one place where you were confident and wrong in a way you would not have remembered.
That second one is the point. It is also the reason the record has to exist before the outcome does — what a Business World Model predicts is only checkable against predictions nobody was able to quietly improve.
Brainis Team
Notes on the company loop — company state, decisions, governed autonomy and verified work — from the people building Brainis and running on it.