Anyone can ask a model why a number moved and get back a confident paragraph. We wanted something an enterprise can actually trust: an engine that writes falsifiable hypotheses, tests each one against live data, and says plainly when it cannot be sure. We named it after Karl Popper.
Every analytics team has sat in this meeting. A number moved. Checkout conversion, activation rate, gross margin, pick your favorite. Eight people are in the room, and within ten minutes there are eight theories about why. The seasoned operator blames the pricing test. The growth lead blames the channel mix. Someone pulls up a dashboard, points at a line, and everyone nods at the line while privately keeping their own theory. The meeting ends. Nothing is settled. The number is still moving.
The instrumentation is rarely the problem. Most companies measure plenty. The problem is interpretation: turning "the number changed" into "here is what changed it, and here is how we know." That gap is where weeks go, and where revenue quietly leaks for an entire quarter before the right answer happens to walk into the right room.
There is a tempting shortcut sitting on everyone’s desk now. Paste the chart into a chat model, ask why conversion dropped, read the fluent paragraph it writes back. It might even be right. The trouble is that you have no way to know whether it is, and neither does the model. It optimized for an answer that sounds correct, not one that is.
A single model call is a remarkable thing. It is also the wrong tool for a question a CFO is about to act on. Asked why a metric fell, a model returns the most probable-sounding story given everything it has read. That story is shaped by the training data and the phrasing of your question, not by what actually happened in your business last Tuesday. It has never seen your checkout funnel this week. It is guessing, gracefully.
The dangerous part is not that models occasionally hallucinate. It is that a lone answer gives you nothing to check it against. It arrives confident and unfalsifiable, which is the exact combination you least want behind a decision. You cannot tell a good answer from a lucky one, because both look identical: a tidy paragraph with no seams.
We named the engine after Karl Popper, the philosopher who spent a career on a deceptively simple question: what separates science from storytelling? His answer was falsifiability.¹ A claim earns the right to be called scientific only if you can state, in advance, the observation that would prove it false.
"All swans are white" is a real claim, because a single black swan ends it. "The dip was caused by market conditions" is not a claim at all, because nothing you could ever observe would disprove it. It explains every outcome equally well, which means it explains nothing. Most of the theories in that Monday meeting are swans of the second kind: unkillable, and therefore useless.
We took the idea literally. Inside Popper, a hypothesis that cannot be refuted does not get tested. It does not earn the compute. Before the engine spends a cent investigating a claim, that claim has to come with its own undoing: the specific evidence that, if found, would settle the matter against it.
“We did not want a model that sounds right. We wanted a system that can be shown to be wrong.”
Start with a brief: a metric, a window of time, and whatever context the person brought. The first thing Popper does is propose. It writes a small set of hypotheses, usually a handful, and each one is a falsifiable claim, not a topic to look into. "Investigate mobile" is not allowed. "The drop is concentrated in mobile checkout, and if web fell by the same amount this is wrong" is.
If you brought a hunch, it goes in first. There is no faster path to a verdict than killing or confirming the obvious suspect, and people are often right about their own systems. Popper sharpens the hunch into a falsifiable form and tests it before anything else, which means the most common outcome of a good guess is that you find out quickly whether it holds.
Then each hypothesis gets its own investigator. They work independently and in parallel, each one pulling the specific data its claim needs and nothing it does not. An investigator returns exactly one of three verdicts: confirmed, refuted, or inconclusive, with a confidence and the evidence it leaned on. They do not talk to each other while they work, which matters more than it sounds: it keeps one investigator’s early guess from quietly biasing the others.
When the round finishes, a reasoner reads the whole grid at once. It sees which hypotheses held, which died, and how strong each call was, and it decides what to do next. It can sharpen a confirmed lead into narrower sub-questions. It can throw out the round and try a genuinely new angle when nothing survived. Or it can decide there is enough and write the verdict.
The detail that keeps this from being a science fair project: it runs on a fixed budget. The number of rounds is bounded. Each investigator has a step limit. Popper is not an open-ended agent that crawls your warehouse until it feels satisfied and hands you a bill. It is a search with a deadline, which is the only kind you can put in front of a customer.
Every round leaves a record behind: the hypothesis, the evidence that would have refuted it, the verdict it reached, the confidence, and the data it touched on the way. Strung together, those records are a ledger of the entire investigation.
Numbers above are illustrative, but the shape is real. This ledger is the artifact that makes a Popper result auditable. You are not asked to trust a conclusion on faith. An analyst can open any line, see what it looked at, and replay the path that led to the answer. The reasoning is the product, not a byproduct of it.
Say cart-to-checkout conversion drops on a Tuesday. Here is the kind of grid Popper builds, and the point is in the second column: every hypothesis arrives with the exact finding that would kill it.
Notice that the refuted row is worth as much as the confirmed one. Learning that it was not the iOS release is not a dead end. It is a team saved from a week of rolling back the wrong thing. Most reasoning systems treat a negative result as a failure to find the answer. We treat it as half the answer.
Let us be fair to the models, because we are heavy users of them. We run frontier models throughout Popper, including for the reasoning itself, and they are extraordinary. But a model is a component, not the product. The hard parts of this system are not the prompt, and they do not fall out of a bigger model or a cleverer wording.
Consider what actually has to be true for a single Popper verdict to be trustworthy:
None of that is a prompt. It is orchestration, retrieval, evaluation, cost control, and a long list of opinionated decisions about what "done" means and what counts as enough evidence. The model is the easy mile. The discipline around it is the marathon, and the marathon is the company.
“The model is the easy part. The discipline around it is the product.”
There is a quieter problem underneath all of this: how do you even know an investigation engine is any good? You cannot grade it on whether the prose reads well, because fluent and wrong is exactly the trap. You have to grade it on whether its refutations actually hold, whether its confirmations survive a second independent look, and whether it reliably knows what it does not know. Building that evaluation is its own discipline, and it is most of the work that never shows up in a demo.
Enterprise buyers are right to be suspicious of a confident black box. We are too. So trust, for us, is not a reassuring voice or a polished summary. It is a set of concrete design choices, and each one buys the reader something specific.
The most counterintuitive choice is the third one. Inconclusive is a first-class verdict in Popper, and reaching it counts as the system working, not failing. A model forced to always produce an answer will, sooner or later, produce a wrong one with total confidence. We would rather it stop and tell you it could not establish the claim. An honest "I could not show this" is worth more to a serious team than a confident guess, every single time.
Back in that Monday meeting. Instead of eight competing theories and a dashboard everyone nods at, you open a short list of hypotheses that were actually tested. The ones that survived. The ones that died, and why. The evidence sitting under each, one click away. You can disagree with the conclusion, which is the point. You can also check it, which is the other point.
And because Popper runs continuously in the background rather than only when summoned, the investigation is often already waiting for you, finished, before you have thought to ask the question. The meeting changes from a debate about what the number means into a review of what the evidence already showed.
A dashboard will tell you that a number changed. It will never tell you why, and it will certainly never tell you when it is unsure. We built Popper because "why" deserves the same rigor the industry has spent two decades pouring into "what." That it tries, every single time, to prove itself wrong is not a hedge or a limitation. It is the whole reason you can believe it when it turns out to be right.
Karl R. Popper, The Logic of Scientific Discovery (Routledge, 1959; originally Logik der Forschung, 1934). The source of the falsifiability criterion.
View sourceKarl R. Popper, Conjectures and Refutations: The Growth of Scientific Knowledge (Routledge, 1963). On knowledge advancing by bold conjecture and rigorous attempts at refutation.
View sourceSalman is the co-founder and CEO of Augmeta. He writes here when he has something he can't shake — usually about why teams keep mistaking dashboards for understanding, and what it would take to fix that.