← SIGNAL
10 min read · Engineering · October 2026

Why did it drop? How we picked a small model to decide when Kai offers a triage

Kai offered a root-cause triage based on the shape of its own plan, not on what the user asked. We tested four ways to read the question instead, on 626 production turns. A small typed-decision model tied Claude Haiku 4.5 at a third of the latency.

M
MaazFounding Engineer, Augmeta

Kai is our analytics assistant. Ask it about a KPI and it answers from a quick look at the data. When you need the real cause, there is a triage: a full investigation across segments, drivers and signals that ends in a root cause you can check.

The hard part is not running a triage. It is knowing when to offer one. Offer it on every turn and people learn to scroll past it. Miss a real why-question and the person who most needs the investigation never sees the door.

So we needed a small, fast model to answer one question on every chat turn: is this person asking why a metric moved? We tested four approaches on 626 real production turns. Jev, TypeSafe's typed-decision model, tied Claude Haiku 4.5 on ranking quality (average precision 0.86 vs 0.84) while running about 3x faster at roughly 1/25th of the cost. An open-source look-alike and a keyword rule were not close.

0.86 vs 0.84
Average precision, Jev vs Claude Haiku 4.5. A statistical tie, at about a third of the latency.

This post covers what we count as a triage question, what was wrong with the old trigger, how we tested, where the test is weaker than we would like, and how Jev runs inside Kai today.

What counts as a triage question

A turn deserves a triage offer when it asks for the cause of an observed movement in a business KPI. Three things must hold:

  • Explanatory intent. The user wants the cause, not the size of the change or a chart of it.
  • A movement that actually happened. Not one the user is imagining or planning for.
  • A business metric as the subject. Something in the org's KPI registry, not a broken page or a slow dashboard.

We wrote that down as a five-label rubric. Only explanatory earns an offer. An explicit_triage request already goes straight to the existing triage tool, so it needs no offer.

User asksLabelWhy
Why did cart-to-checkout fall on the 5th of August?explanatoryAsks for the cause of an observed drop
What drove the spike in desktop visits last week?explanatoryCausal ask without the word "why"
Is the conversion dip yesterday because of the app release?explanatoryTests a causal hypothesis
Why? (after Kai reported a 12% drop in AOV)explanatoryFollow-up; the KPI comes from the previous turn
We raised shipping fees last week and conversion dropped. Why would that happen?explanatoryThe drop is real; the user wants its cause
Would conversion drop if we raised shipping fees?hypotheticalNo movement has happened yet
How much did cart-to-checkout fall on the 5th of August?descriptiveAsks for size, not cause
Run a triage on conversion rate for yesterdayexplicit_triageHandled by the existing triage tool
Why is the KPI chart not loading?otherProduct issue, not a metric
Only explanatory turns get an offer. The two shipping-fee rows share most of their words and land on opposite sides of the line.

The label depends on what is being asked, not on any single word. "Why would conversion drop if we raised fees?" is genuinely ambiguous on its own: it could be a what-if, or it could be someone who already raised fees and is looking at the damage. Telling those apart, from the message or from the chat around it, is exactly the job the detector has to do.

Why not just write a rule?

A teammate raised the obvious objection on an early draft: if a triage question is "why", plus a KPI, plus a real movement, isn't that just a boolean?

python
offer_triage = kpi_in_registry and kpi_has_moved and "why" in message.lower()

It is a fair statement of the idea, and each part breaks on real traffic:

  • "Why" is not necessary. "What drove the spike in desktop visits?" has no why. Of 126 briefs from triages users actually started, 28 contain no question at all, just statements like "mobile web cart to checkout has been trending downwards".
  • "Why" is not sufficient. "Why is the KPI chart not loading?" has the word and a metric-ish subject. "Why would conversion drop if we raised shipping fees?" has the word and a real registry KPI, and still describes a drop that never happened.
  • "Has moved" is not something we can look up yet. The check runs the moment the message arrives, before Kai has fetched any data. At that point "has the KPI moved" means "does the user say it moved", which is a reading problem, not a lookup.
  • "In the registry" is fuzzy too. Users write "CTC", "checkout conversion" and "cart-to-checkout" for the same KPI. And a KPI named after a page reads like an outage: "cart-to-checkout is down" is a metric falling, while "checkout is down" is a broken page.

We still tested a keyword rule as the floor. It reached 0.45 precision and 0.53 recall, meaning it missed nearly half the real triage questions and was wrong more often than right when it fired.

What was wrong with the old trigger

Kai already showed a "Go deeper with a triage" card under some answers. The trouble was what decided it. The card was keyed to the shape of the plan Kai ran: if the plan fetched a KPI, a card appeared at the end of the answer. It never looked at what the user had asked.

User asksOld trigger didWhat should happen
"What was cart-to-checkout on Aug 5?"Offered a triage. The plan fetched a KPI, so a lookup looked like an investigation.No offer. They asked for a number, not a cause.
"Why did conversion drop yesterday?", answered inlineNo offer. Kai answered without a plan, so the card code never ran.Offer a triage.
"Why did cart-to-checkout fall on Aug 5?"Opened on whichever KPI the plan fetched first, sometimes a driver metric, over the plan's 14- or 30-day context window.Open on cart-to-checkout, for Aug 5.
Any why-questionOffered only after Kai's full answer: 3.5 minutes in one staging test.Offer before Kai starts working.
Every failure traces back to one mistake: deciding from what Kai did, not from what the user asked.

Each of these costs the user something different:

  • Offers on lookups are noise. A card under "what was the number on Aug 5?" teaches people the card means nothing, so they ignore it when it matters.
  • Missing inline answers hides the triage from the people asking for it. Whether Kai builds a plan is our implementation detail: for simple questions it answers from a number or two without one. A user who asks why conversion dropped wants the cause either way, and a quick inline answer is a glance at the data, not an investigation of segments and drivers. The offer should depend on the question, not on which path Kai took internally.
  • The wrong KPI or window investigates the wrong thing. Either the triage runs on a driver metric over a month, or the user has to notice and fix it by hand in the picker.
  • A late offer wastes the wait. By the time it appears, the user has sat through minutes of analysis that the triage would redo more thoroughly.

There was a quieter bug underneath: the browser and Kai resolved metric names separately and could disagree about which KPI the card meant. The fix starts with the first problem, though. Trigger on the user's intent, not on the plan.

How we tested

We pulled 626 turns from production chat traces and ran every candidate detector on the same set:

  • 400 random Kai turns. The realistic mix; 46 of them were positives.
  • 100 turns pre-filtered for why-style words. A harder slice, dense with near-misses.
  • 126 briefs from triages users actually started. What people wrote when they clearly wanted a triage.

Reference labels, and their weakness. Claude Opus 5 applied the rubric above to every turn. That choice tilts the test toward Haiku: Haiku comes from the same model family and was prompted with the same rubric, so where the two agree the test cannot fully separate agreement with the labeller from being right. A reviewer called this close to cheating, and that is fair. We kept it for this round for two reasons. The tilt runs toward Haiku, the model Jev was up against, so it cannot explain Jev keeping pace. And the 126 user-started triages carry a label no model wrote: the user really did start a triage. We report that slice on its own below. Relabelling with a model from outside both families is the next step.

Where the scores come from. Every detector has to output a score so we can rank turns and pick a cut-off. Jev and Laya return a probability for each answer they are given; we used the probability of "yes". Haiku was asked to return a confidence between 0 and 1 alongside its label. The regex either matches or it does not.

Tuning. No model was fine-tuned for the comparison, Haiku included. "Tuning" here means one thing: on half of the random turns we picked each detector's threshold to hit 0.90 precision, then scored it at that threshold on the other half. "At the bar" below means at that tuned threshold.

Metrics. Precision and recall at the bar, plus average precision: how well a detector ranks true triage questions above everything else, across all thresholds. Average precision is the fairest single number here, because it does not depend on where one threshold happens to land.

The contenders:

  • Jev (TypeSafe, pinned to jev-1.13.0) in three setups: the full rubric plus the previous turn, the rubric on the message alone, and a single yes/no question plus the previous turn.
  • Claude Haiku 4.5, with and without the previous turn.
  • Laya, an open-source model built in Jev's image that runs locally in about 30 ms, in three setups: one yes/no, three yes/nos multiplied, and a five-label pick.
  • A keyword regex as the floor.

Results

Jev with the rubric and the previous turn came out on top, level with Haiku and far ahead of everything that runs locally.

DetectorAverage precision
Jev, rubric + previous turn (chosen)0.86
Haiku 4.5, with previous turn0.84
Jev, rubric, message only0.81
Jev, single yes/no + previous turn0.81
Haiku 4.5, message only0.77
Laya, single yes/no (best wording)0.52
Laya, three yes/nos multiplied0.39
Keyword regex0.33
Laya, five-label pick0.21
Experiment of 2026-09-23: 9 detector setups, 626 production turns. 1.0 is a perfect ranking.

The 95% confidence interval on the difference between Jev and Haiku runs from -0.10 to +0.09: a tie, not a win. Jev beat the best Laya setup by +0.34 (95% CI +0.20 to +0.48), which is not close.

At the tuned threshold, with speed included:

DetectorPrecision at the barRecall at the barLatency (median / p95)
Jev, rubric + previous turn (chosen)0.850.770.35 s / 0.41 s
Jev, rubric, message only0.830.750.35 s / 0.43 s
Jev, single yes/no + previous turn0.810.710.35 s / 0.42 s
Haiku 4.5, with previous turn0.920.741.06 s / 2.03 s
Haiku 4.5, message only0.850.641.03 s / 1.93 s
Laya, single yes/no (best wording)0.390.680.16 s for 5 questions
Laya, three yes/nos multiplied0.360.59same call
Laya, five-label pick0.130.97same call
Keyword regex0.450.53none

Jev vs Claude Haiku 4.5

Haiku edges Jev on precision at the bar (0.92 vs 0.85). That edge is measured against Opus applying Haiku's own rubric, so part of it is agreement with the labeller rather than correctness. Reading Jev's highest-scored "false positives" supports that: several are real triage questions the labels missed, such as "create a one-pager narrative explaining the downward trend".

On everything else Jev wins: 350 ms median against about 1 s, 410 ms p95 against 2 s, and about $0.00004 per call against $0.001. Jev also returns probabilities and typed choices, never free text, so there is nothing to parse.

Jev vs Laya

Laya is modelled on Jev, so we hoped it would behave like Jev for free. It did not. Zero-shot, its scores bunched between 0.3 and 0.7, so no threshold meant much; Jev's scores spread from 0.01 to 0.99. Fine-tuning Laya on our data lifted it to 0.65 average precision, still well short. "Built like Jev" turned out to be no evidence it would work on our task. Only measuring told us that.

Jev vs a keyword rule

The regex failed in both directions, as the rule section predicted: "What drove the spike" has no why, and "why is the chart not loading" is not a KPI question.

The check no model labelled

The 126 user-started triages are our cleanest evidence, because the label is an action a person took, not a model's opinion. They tell a sobering story: every method caught only 60 to 72% of them. 28 of those briefs contain no why-question at all, just statements like "mobile web cart to checkout has been trending downwards". Our rubric says statements like that do not need a triage. Users clearly disagreed.

That is a definition problem, not a model problem, and no amount of model-shopping fixes it. It also means this slice does not separate Jev from Haiku; it shows the ceiling they share. Separating them cleanly is what the independent relabel is for.

How much of the chat should the detector see?

Another review question: wouldn't more turns of context simply give better results? The experiment says yes, and production has since gone further than the experiment did.

ContextJevHaiku 4.5
Current message only0.810.77
Plus the previous turn0.860.84
Average precision on 626 turns. One turn of context was worth more than the gap between the two models.

Jev's remaining misses were mostly short follow-ups like "okay do it" or "yes, go ahead with the next steps". The experiment passed only the first 300 characters of the previous turn, which is not enough to know what "it" is. QA on staging hit the sharpest version of this: Kai answered a lookup, ended with "Would you like me to dig into why?", the user said "yes", and no triage was offered. The offer sat at the end of Kai's reply, and the first 300 characters could never see it.

So Jev now gets the whole chat. The current message goes as before, and every earlier user message and Kai reply goes along as conversation, oldest first. Kai's replies are cut to their last 600 characters, because what matters is how they end. A cap of 60 messages only guards against runaway threads. The rubric limits what the history can do: it is used to resolve references like "why?", "yes" and "go ahead", and an earlier why-question does not make an unrelated message count.

Measured on the same turns, average precision went from 0.82 to 0.87. On the 279 follow-up messages, precision rose from 0.68 to 0.78 and recall from 0.59 to 0.72. One caveat: the dataset kept only the previous turn for almost every trace, so this measures the new format with one turn of history, not the full window. The longer histories are checked live instead. The QA thread's "yes" now scores 0.94 and "go ahead" 0.93, while "yes" to "want me to export a CSV?" stays at 0.08 and "make the chart a bar chart" after a why-question stays at 0.04.

How Jev runs inside Kai

The check is its own step at the start of every chat turn, before Kai plans anything. If Jev says yes, Kai pauses with two choices, "Start a triage" and "Continue with Kai", and names the window it read ("Start a triage for 14 to 20 Sep?"). If Jev says no, or anything fails, the turn runs exactly as it would have without the check. It adds about 250 ms on a warm connection and shows nothing on screen while it runs.

Each check sends two requests in parallel:

  • Detection. A single needs_triage yes/no, with the rubric passed in as the labelling policy. Jev returns a probability, and the offer needs at least 0.72. The threshold was re-tuned for the full-chat setup, where it gives 0.84 precision and 0.83 recall on the random turns. The older labels still score some page-named metric drops ("cart-to-checkout down") as outages, so true precision is somewhat higher.
  • The date window. About 20 typed choices: kind of window, month, day, year, weekday, lookback length, named period. Jev never writes a date. Our code does the calendar maths and checks the result, and a window whose weakest part scored under 0.60 confidence becomes no window rather than a wrong one.

After the user says yes. The KPI is Kai's own pick: the card runs Kai's agent, with the same prompt and the same KPI lookup it would use before fetching data, so the triage opens on exactly the KPI Kai would have analysed. That also retires the second resolver in the browser. Where Kai would ask the user which KPI they mean, the picker asks instead. Users who type "sure" or "triage it" instead of clicking get the triage too: Jev reads a typed reply as a choice between the two buttons.

Limits and failure handling. The model is pinned to jev-1.13.0, never latest, because the threshold is tuned to one version. Timeouts are 1.5 s for detection and 2.5 s for the window. Any error or timeout is logged and means no offer. A slow window request drops only the window.

Kill switches and audit. Environment flags turn the check off or move the threshold without a deploy. Every verdict, a "no" included, is recorded on the turn and its trace: score, threshold, window, window confidence, latency and Jev's reported version, so we can audit precision from production.

Why typed choices beat free text for dates

Letting a model write "2026-08-05" means parsing and trusting a string. Asking Jev to pick parts means code can validate every piece: an impossible date becomes none, a missing year means the most recent past one, and everything is clamped to the last complete day.

User wrote (today = Wed 2026-09-23)Result
"on the 5th of August"single day, 2026-08-05
"yesterday"single day, 2026-09-22
"Aug 1-7"range, 2026-08-01 to 2026-08-07
"past 14 days"lookback of 14 days
"today"2026-09-22, the last complete day, with a "today is still in progress" note
"FW37", "early Feb"none; the phrase goes into the triage brief
Fiscal weeks are the known gap: until an org has a fiscal calendar configured, they resolve to none and the user picks dates in the picker.

What we learned, and what's next

A small typed-decision model matched a general LLM on this task, and won on speed, cost and checkable output. A few lessons stood out:

  • Define the question before shopping for a model. Most of the hard cases were arguments about the rubric, not about the models.
  • Measure the look-alike. Laya was built in Jev's image and still failed zero-shot. Architecture is not evidence.
  • Watch who wrote the labels. Our labeller shared a family with one contender. Keeping a slice labelled by real user behaviour was what let us check the results anyway.
  • Context beats model choice. Adding the previous turn moved average precision more than switching between the two best models did.
  • Score spread matters. A detector whose scores bunch in the middle gives you a threshold in name only.
  • Typed output is easier to trust. Choices you can validate beat strings you have to parse, especially for dates.

And next:

  • Independent labels. Relabel the full set with a model from outside the Anthropic family, review the disagreements by hand, and rerun the comparison.
  • Widen the definition. About a third of real triages started from a statement, not a question. The rubric should count "worrying movement" statements as explanatory.
  • Evaluate the full window. Store full chat histories in the dataset, so the whole-chat setup is measured rather than spot-checked.
  • Fiscal calendars. So "FW37" resolves to real dates instead of falling back to the picker.

Keep reading.

All posts