Kai offered a root-cause triage based on the shape of its own plan, not on what the user asked. We tested four ways to read the question instead, on 626 production turns. A small typed-decision model tied Claude Haiku 4.5 at a third of the latency.
Kai is our analytics assistant. Ask it about a KPI and it answers from a quick look at the data. When you need the real cause, there is a triage: a full investigation across segments, drivers and signals that ends in a root cause you can check.
The hard part is not running a triage. It is knowing when to offer one. Offer it on every turn and people learn to scroll past it. Miss a real why-question and the person who most needs the investigation never sees the door.
So we needed a small, fast model to answer one question on every chat turn: is this person asking why a metric moved? We tested four approaches on 626 real production turns. Jev, TypeSafe's typed-decision model, tied Claude Haiku 4.5 on ranking quality (average precision 0.86 vs 0.84) while running about 3x faster at roughly 1/25th of the cost. An open-source look-alike and a keyword rule were not close.
This post covers what we count as a triage question, what was wrong with the old trigger, how we tested, where the test is weaker than we would like, and how Jev runs inside Kai today.
A turn deserves a triage offer when it asks for the cause of an observed movement in a business KPI. Three things must hold:
We wrote that down as a five-label rubric. Only explanatory earns an offer. An explicit_triage request already goes straight to the existing triage tool, so it needs no offer.
The label depends on what is being asked, not on any single word. "Why would conversion drop if we raised fees?" is genuinely ambiguous on its own: it could be a what-if, or it could be someone who already raised fees and is looking at the damage. Telling those apart, from the message or from the chat around it, is exactly the job the detector has to do.
A teammate raised the obvious objection on an early draft: if a triage question is "why", plus a KPI, plus a real movement, isn't that just a boolean?
It is a fair statement of the idea, and each part breaks on real traffic:
We still tested a keyword rule as the floor. It reached 0.45 precision and 0.53 recall, meaning it missed nearly half the real triage questions and was wrong more often than right when it fired.
Kai already showed a "Go deeper with a triage" card under some answers. The trouble was what decided it. The card was keyed to the shape of the plan Kai ran: if the plan fetched a KPI, a card appeared at the end of the answer. It never looked at what the user had asked.
Each of these costs the user something different:
There was a quieter bug underneath: the browser and Kai resolved metric names separately and could disagree about which KPI the card meant. The fix starts with the first problem, though. Trigger on the user's intent, not on the plan.
We pulled 626 turns from production chat traces and ran every candidate detector on the same set:
Reference labels, and their weakness. Claude Opus 5 applied the rubric above to every turn. That choice tilts the test toward Haiku: Haiku comes from the same model family and was prompted with the same rubric, so where the two agree the test cannot fully separate agreement with the labeller from being right. A reviewer called this close to cheating, and that is fair. We kept it for this round for two reasons. The tilt runs toward Haiku, the model Jev was up against, so it cannot explain Jev keeping pace. And the 126 user-started triages carry a label no model wrote: the user really did start a triage. We report that slice on its own below. Relabelling with a model from outside both families is the next step.
Where the scores come from. Every detector has to output a score so we can rank turns and pick a cut-off. Jev and Laya return a probability for each answer they are given; we used the probability of "yes". Haiku was asked to return a confidence between 0 and 1 alongside its label. The regex either matches or it does not.
Tuning. No model was fine-tuned for the comparison, Haiku included. "Tuning" here means one thing: on half of the random turns we picked each detector's threshold to hit 0.90 precision, then scored it at that threshold on the other half. "At the bar" below means at that tuned threshold.
Metrics. Precision and recall at the bar, plus average precision: how well a detector ranks true triage questions above everything else, across all thresholds. Average precision is the fairest single number here, because it does not depend on where one threshold happens to land.
The contenders:
Jev with the rubric and the previous turn came out on top, level with Haiku and far ahead of everything that runs locally.
The 95% confidence interval on the difference between Jev and Haiku runs from -0.10 to +0.09: a tie, not a win. Jev beat the best Laya setup by +0.34 (95% CI +0.20 to +0.48), which is not close.
At the tuned threshold, with speed included:
Haiku edges Jev on precision at the bar (0.92 vs 0.85). That edge is measured against Opus applying Haiku's own rubric, so part of it is agreement with the labeller rather than correctness. Reading Jev's highest-scored "false positives" supports that: several are real triage questions the labels missed, such as "create a one-pager narrative explaining the downward trend".
On everything else Jev wins: 350 ms median against about 1 s, 410 ms p95 against 2 s, and about $0.00004 per call against $0.001. Jev also returns probabilities and typed choices, never free text, so there is nothing to parse.
Laya is modelled on Jev, so we hoped it would behave like Jev for free. It did not. Zero-shot, its scores bunched between 0.3 and 0.7, so no threshold meant much; Jev's scores spread from 0.01 to 0.99. Fine-tuning Laya on our data lifted it to 0.65 average precision, still well short. "Built like Jev" turned out to be no evidence it would work on our task. Only measuring told us that.
The regex failed in both directions, as the rule section predicted: "What drove the spike" has no why, and "why is the chart not loading" is not a KPI question.
The 126 user-started triages are our cleanest evidence, because the label is an action a person took, not a model's opinion. They tell a sobering story: every method caught only 60 to 72% of them. 28 of those briefs contain no why-question at all, just statements like "mobile web cart to checkout has been trending downwards". Our rubric says statements like that do not need a triage. Users clearly disagreed.
That is a definition problem, not a model problem, and no amount of model-shopping fixes it. It also means this slice does not separate Jev from Haiku; it shows the ceiling they share. Separating them cleanly is what the independent relabel is for.
Another review question: wouldn't more turns of context simply give better results? The experiment says yes, and production has since gone further than the experiment did.
Jev's remaining misses were mostly short follow-ups like "okay do it" or "yes, go ahead with the next steps". The experiment passed only the first 300 characters of the previous turn, which is not enough to know what "it" is. QA on staging hit the sharpest version of this: Kai answered a lookup, ended with "Would you like me to dig into why?", the user said "yes", and no triage was offered. The offer sat at the end of Kai's reply, and the first 300 characters could never see it.
So Jev now gets the whole chat. The current message goes as before, and every earlier user message and Kai reply goes along as conversation, oldest first. Kai's replies are cut to their last 600 characters, because what matters is how they end. A cap of 60 messages only guards against runaway threads. The rubric limits what the history can do: it is used to resolve references like "why?", "yes" and "go ahead", and an earlier why-question does not make an unrelated message count.
Measured on the same turns, average precision went from 0.82 to 0.87. On the 279 follow-up messages, precision rose from 0.68 to 0.78 and recall from 0.59 to 0.72. One caveat: the dataset kept only the previous turn for almost every trace, so this measures the new format with one turn of history, not the full window. The longer histories are checked live instead. The QA thread's "yes" now scores 0.94 and "go ahead" 0.93, while "yes" to "want me to export a CSV?" stays at 0.08 and "make the chart a bar chart" after a why-question stays at 0.04.
The check is its own step at the start of every chat turn, before Kai plans anything. If Jev says yes, Kai pauses with two choices, "Start a triage" and "Continue with Kai", and names the window it read ("Start a triage for 14 to 20 Sep?"). If Jev says no, or anything fails, the turn runs exactly as it would have without the check. It adds about 250 ms on a warm connection and shows nothing on screen while it runs.
Each check sends two requests in parallel:
needs_triage yes/no, with the rubric passed in as the labelling policy. Jev returns a probability, and the offer needs at least 0.72. The threshold was re-tuned for the full-chat setup, where it gives 0.84 precision and 0.83 recall on the random turns. The older labels still score some page-named metric drops ("cart-to-checkout down") as outages, so true precision is somewhat higher.After the user says yes. The KPI is Kai's own pick: the card runs Kai's agent, with the same prompt and the same KPI lookup it would use before fetching data, so the triage opens on exactly the KPI Kai would have analysed. That also retires the second resolver in the browser. Where Kai would ask the user which KPI they mean, the picker asks instead. Users who type "sure" or "triage it" instead of clicking get the triage too: Jev reads a typed reply as a choice between the two buttons.
Limits and failure handling. The model is pinned to jev-1.13.0, never latest, because the threshold is tuned to one version. Timeouts are 1.5 s for detection and 2.5 s for the window. Any error or timeout is logged and means no offer. A slow window request drops only the window.
Kill switches and audit. Environment flags turn the check off or move the threshold without a deploy. Every verdict, a "no" included, is recorded on the turn and its trace: score, threshold, window, window confidence, latency and Jev's reported version, so we can audit precision from production.
Letting a model write "2026-08-05" means parsing and trusting a string. Asking Jev to pick parts means code can validate every piece: an impossible date becomes none, a missing year means the most recent past one, and everything is clamped to the last complete day.
A small typed-decision model matched a general LLM on this task, and won on speed, cost and checkable output. A few lessons stood out:
And next: