"World model" is a phrase doing a lot of unearned work at the moment, so it is worth being concrete about what one predicts and what it does not.
A Business World Model does not predict the future. It predicts the consequence of a decision, given a state, expressed as a range, with a recorded basis. That is a much narrower thing, and the narrowness is what makes it useful.
What you’ll learn
- What a world model is a model of, precisely
- The five-layer stack, and which layers are rented
- The four question shapes it answers
- What it declines to answer, and why abstention is a feature
A model of decisions, not of weather
The everyday intuition for prediction comes from forecasting: a system observes conditions and tells you what will happen. Business prediction is not that shape, because in business the thing you want to know is conditional.
You are not asking "what will revenue be". You are asking "what happens to revenue if we move the two engineers off the integration and onto the onboarding flow, given that the pipeline looks like this and the team looks like that". The answer to the second question is actionable. The answer to the first is a horoscope with a spreadsheet.
So the object being modeled is the mapping from state and decision to outcome. Everything else follows from that.
The stack, and the honest part
The Business World Model is a stack, not a single network. Rented frontier models sit on top as a replaceable layer. The prediction, decision-quality, calibration, and simulation layers underneath are ours.
That distinction matters commercially and it matters for how you should read any claim in this space, ours included.
The top layer is language and reasoning, and it is rented. It changes every few months, it is available to everyone, and treating it as a moat would be a mistake. What sits underneath is the part that has to be built: the representation of company state, the machinery that turns a proposed decision into a distribution over outcomes, the grading loop that will score those distributions against what happened, and the simulation that lets a path be explored before it is taken.
We are not pretraining a frontier language model, and it would be noise to claim otherwise. The defensible asset is the dataset shape: longitudinal, structure-only sequences of state, decision, governed execution, independent verification, and measured outcome, across many companies. That data does not exist in text about business. It only exists where the loop is actually running, which is why it is hard to acquire and why acquiring it is the whole strategy.
Important: A model trained on writing about business has read every case study. It has never seen a decision it made get graded. Those are different kinds of knowledge, and only one of them improves.
The four question shapes
In practice the questions a business world model is asked collapse into four shapes.
Consequence. If we take this action, what happens to the outcomes we care about, and by when? Returns a range per outcome, with the assumptions it leaned on made explicit.
Comparison. Of these three paths, which dominates, and under what conditions does the ranking flip? Often the flip condition is more useful than the ranking, because it tells you what to watch.
Sensitivity. Which single assumption, if wrong, would most change this recommendation? This is the question that finds the load-bearing belief nobody had examined.
Detection. Given the state as it stands, what is likely to happen that nobody has asked about? The unprompted question — the one a dashboard cannot ask on your behalf.
The strategy layer that turns these answers into a chosen path is described on the Strategy Engine; Brainis models alternative futures before recommending one rather than proposing the first plausible plan.
Every prediction is registered
Registration ships today. The prediction is written down with its range, its assumptions, and the model version that produced it — before the decision is taken, not reconstructed afterwards — and it is sealed, so it cannot be quietly revised once the answer is known.
The mechanical commitment on top of that is simple to state and unpleasant to implement: grading should be automatic, comparing the outcome against what was predicted and recording the result whether or not anyone is looking. That half is not built yet. Scoring a prediction against its outcome is a human step today, which means we publish no calibration figure and you should not infer one.
We are stating it as a commitment rather than a feature because the distinction matters to the argument: without the automatic half, calibration becomes a project somebody is supposed to run, and projects somebody is supposed to run do not survive a busy quarter. That is precisely why it is the next thing being built rather than something we are content to leave to discipline. Writing predictions down before acting is the same discipline described from the operator's side.
What it declines to answer
The most important behaviour is refusal.
Ranges stay honest-wide until calibration evidence justifies narrowing them, and a prediction that cannot be made honestly is an abstention rather than a guess. A model with thin evidence for a question should return a band so wide it is visibly unhelpful, or return nothing at all with a reason.
This is unnatural for language models, which are built to produce fluent output and will produce it regardless. It is also the property that decides whether a prediction layer is worth having. A system that always answers teaches you to stop reading its confidence, and once you stop reading its confidence you are back to guessing with extra steps.
The corollary is that a wide band is a real answer rather than a failure to answer, which is the case ranges beat certainty makes at length.
Three consequences follow, and they are worth knowing before you evaluate anything in this category:
What this buys a founder
You do not get certainty. Nothing in this category gets you certainty, and the vendors implying otherwise are describing an outcome no architecture produces.
What you get is a decision surface where the assumptions are explicit, the ranges are honest, the sensitivity is named, and the record of how well the model reads your business grows every week whether or not anyone maintains it. That is a better basis for a hard call than a confident number with no history behind it.
The architecture, layer by layer, is on the Business World Model page.
Brainis Team
Notes on the company loop — company state, decisions, governed autonomy and verified work — from the people building Brainis and running on it.