Or press ESC to close.

Your Safer Model Just Failed the Suite: Why pass rate is the wrong top-line metric for AI features

Oct 4th 2026 17 min read
medium
javascriptES6
ai/ml
reporting
strategy

A team swaps a newer model into an AI-powered feature, reruns the regression suite, and watches the pass rate fall. The instinctive reading is a regression: block the release, investigate, roll back. But a pass rate only counts the test cases that went red, and it says nothing about what those cases would cost in production. This post uses a marketplace listing reviewer as a running example to show how the model that scores worse on the dashboard can be the safer one, and why the number worth reporting for an AI feature is built from individual decisions priced by their consequences.

The Release That Got Safer and Went Red

Picture a second-hand marketplace that uses an LLM to review every new listing against its policy. For each listing the model returns one of three verdicts: approve, reject, or escalate to a human reviewer. The regression suite holds 120 listings. Each one is checked against four rules (prohibited items, recalled products, misleading claims, and off-platform contact details), and a test case passes only if all four calls match the expected ones. The release gate is simple: the suite must stay at or above an 80% pass rate.

The current model passes 101 of the 120 cases, which is 84%. The team upgrades to a newer model that was tuned to be more cautious and reruns the suite. This time 87 cases pass, which is 72.5%. The gate fails, CI turns red, and the release is blocked. Nobody is being unreasonable here. The count of red cases went from 19 to 33, and that looks exactly like what a regression looks like.

Now look at what each model actually did with the listings that break policy. The current model waved nine violations through to the live site, including the kind of listing nobody wants to explain afterwards, such as a recalled child car seat. The newer model waved through two. On the thing the feature exists to prevent, it is considerably better. The extra red cases are mostly something else entirely: listings where the model declined to pick a side and sent the item to a human, because the suite expected an approve or a reject and got an escalate instead.

Those cases are red, and so is the car seat that went live. Pass/fail records only that the outcome differed from the expected one, not in which direction or at what cost. Both numbers above are correct, and they describe the same runs. The rest of this post pulls them apart, prices each kind of miss, and reconciles the red cases back to the decisions behind them.

The Running Example: A Listing Reviewer

The feature itself is deliberately ordinary. A marketplace receives far more new listings than people can read, so a model does the first pass. It is given the listing (title, description, category, and price) along with the marketplace's policy, and it has to answer one question for each rule in that policy: is the listing fine, does it break this rule, or is it unclear enough that a person should look? The answer comes back as a verdict and a short reason. Nothing about this depends on a particular vendor or model.

The important detail is that the model makes one call per rule, not one call per listing. Here approve means the listing is fine on that rule, reject means it violates the rule, and escalate means the model is passing the decision to a human. What happens to the listing as a whole follows from its four calls: it goes live only if every rule approves it, any reject keeps it off the site, and anything left over lands in a reviewer's queue. That structure is why a suite of 120 listings is really a suite of 480 individual decisions, a fact that becomes central once the pass rate starts to disagree with itself. Each test case in the suite is a listing plus the expected verdict for each rule:

                
{
  id: 'listing-047',
  title: 'Antique hunting knife, 18 cm blade, collector piece',
  category: 'collectibles',
  price: 85,
  expected: {
    prohibited_items: 'escalate',
    recalled_products: 'approve',
    misleading_claims: 'approve',
    off_platform_contact: 'approve',
  },
}
                

Three of the four expected values are unremarkable. The fourth is the interesting one: the correct answer for prohibited_items is escalate. The policy allows antique blades as collectibles but restricts them by length and by region, and a title and a price are not enough to settle which side of that line this listing sits on. The right behavior is neither to approve it nor to reject it, but to ask a person. The expected value is therefore three-valued, exactly like the model's output, and that detail matters later. It means the model can miss in more ways than giving the wrong answer, and one of those ways is being sure when it should have been unsure.

What a Binary Pass/Fail Throws Away

A test assertion asks one question: does the actual value equal the expected one? If it does, the case is green, and if it does not, it is red. For deterministic code this is a perfectly good design, because a mismatch usually means a bug, and bugs are things you want to hear about regardless of their flavor. The assumption underneath is that the ways of being wrong are roughly interchangeable. For a model that chooses between three verdicts, that assumption does not hold. Here is what three red lines from the new model's run look like in a typical CI report:

FAIL  listing-012  recalled_products  expected: reject   actual: approve
FAIL  listing-031  misleading_claims  expected: approve  actual: reject
FAIL  listing-058  prohibited_items   expected: approve  actual: escalate

In the report these are three identical entries. Each one is a single failure, each one counts once against the pass rate, and each one is exactly as red as the others. In the world the marketplace operates in, they are three completely different events. The first is a recalled product that is now live and could end up in someone's home. The second is an honest seller whose perfectly good listing was blocked. The third is a listing that a reviewer will glance at and approve. The consequences differ by orders of magnitude, and the report cannot show it, because "expected is not equal to actual" is all it knows.

That single bit of information discards three separate things, each of which a release decision actually depends on:

The pass rate averages over failures that were never comparable, and that is how a release gets blocked for the wrong reason. A model tuned to be more cautious trades expensive mistakes for cheap ones: fewer violations wave through, and more ambiguous listings go to a human. For the business that is often exactly the trade worth making. For an assertion it is a loss, because the cheap outcome is still a mismatch and there are more of them. A gate built on the pass rate cannot tell this kind of improvement from a regression, and fixing that takes more than a different threshold. It takes a different thing to count.

The Unit Is the Decision, Not the Test Case

The pass rate has a second problem, and it sits one level above the assertion. The metric counts test cases, and a test case in this suite is a bundle of four decisions. A case turns red when any one of those decisions is a mismatch, and it stays exactly as red if two more go wrong. Here are two red cases from the new model's run, opened up to show the decisions inside them:

listing-071  FAIL
  ok    prohibited_items      expected: approve  actual: approve
  MISS  recalled_products     expected: reject   actual: approve
  MISS  misleading_claims     expected: approve  actual: escalate
  MISS  off_platform_contact  expected: reject   actual: escalate

listing-103  FAIL
  ok    prohibited_items      expected: approve  actual: approve
  ok    recalled_products     expected: approve  actual: approve
  MISS  misleading_claims     expected: approve  actual: escalate
  ok    off_platform_contact  expected: approve  actual: approve

To the pass rate these are the same thing: two red cases, one failure each. Inside, they are nothing alike. Listing 103 contains a single hedge on a claim a human will wave through in a minute. Listing 071 contains a recalled product that the model approved, plus two more hedges alongside it. The first is a trivial annoyance and the second puts a dangerous item on the site, and both are one unit of red.

The count is misleading in both directions. A red case is not wrong everywhere, since listing 071 still got one of its four decisions right, so counting it as a whole failure overstates what went wrong. But it can also hold several misses, so counting it as one understates how many mistakes were made. Red means "at least one decision disagreed", and nothing more. That is why the number of red cases and the number of wrong decisions will never match, and why the gap is not an error in either count.

The decision is the right unit because it is the thing the system actually does. Nothing in production happens to a test case. A recalled product goes live, a seller is blocked, a reviewer picks up a ticket. Each of those is the consequence of one call the model made on one rule, and each has its own price. If the aim is to measure what the feature costs, the decision is the smallest piece that still carries a cost, and the test case is only the container it arrived in.

Changing the unit does not mean writing new tests. The suite already compares expected to actual for every rule of every listing, then folds the results into a single boolean per case and throws the detail away. Keep it instead. Record one result per decision, with the case id, the rule, the expected verdict, and the actual verdict, and derive the case-level pass or fail from those records if you still want it. The case stays the unit of authoring and the decision becomes the unit of measurement.

Five Outcomes, Five Prices

Once the decision is the unit, every decision can be described with three things: its outcome, which is the category it falls into, what happened, which is the plain description of the event, and what it cost, which is the consequence of that event for the business. For the listing reviewer there are five outcomes:

Outcome What happened What it cost
Agreement The model matched the expected call Nothing
False negative A violation was waved through A takedown, an internal escalation, and regulatory and reputational exposure. Roughly $400
False positive A compliant listing was rejected A support ticket and a seller who may not come back. Roughly $15
Inconclusive The model declined to decide and escalated About four minutes of reviewer time. Roughly $4
Mirror case The model decided confidently where a hedge was the right answer Depends on the direction of the call, and nothing if it happened to land right. Roughly $15 in this suite

A false negative is the expensive one, because the damage has already reached the site by the time anyone notices. A false positive costs a seller's patience and a support ticket. An inconclusive decision is the cheapest way to be unhelpful. Note that escalating is also a correct answer when the expected verdict is escalate. That counts as an agreement and costs nothing.

The fifth row is the mirror of the inconclusive one, and it is easy to overlook because it never looks like a failure. Inconclusive is the model hedging when the expected answer was a decision. The mirror case is the model deciding when the expected answer was a hedge, like approving the antique hunting knife from the fixture with full confidence instead of passing it to a person. If the guess happens to be right, nothing bad follows, but the review that should have taken place did not, and the system was lucky rather than careful. If the guess is wrong, the price is that of a false negative or a false positive, depending on the direction, with the difference that no one was ever asked. In the new run all three mirror cases were confident rejections of borderline listings, so each is priced like a false positive. A confident approval of a borderline listing would be priced like a false negative, and it is the one worth watching most closely.

The dollar figures are placeholders, and every team would calibrate its own from support costs, reviewer time, and past incidents. What carries over is the shape: a false negative costs roughly a hundred times an inconclusive decision and more than twenty times a false positive, so the outcomes cannot be added up as if each were worth one.

One property makes the rest of the post possible. Every decision belongs to exactly one of these five outcomes, with no overlap and nothing left out, so the outcome counts must add up to the total number of decisions. That is what will let two numbers that look contradictory be shown to be the same data. First, the buckets need to be assigned by code instead of by eye.

Classifying a Decision in Code

The classifier is a small pure function. It takes the expected verdict and the actual verdict for one decision and returns the name of one outcome. It makes no model call and has no state, which is what makes it trustworthy as a measuring instrument: if the model is the thing being measured, the code that does the measuring should be boring.

                
const VERDICTS = ['approve', 'reject', 'escalate'];

export const OUTCOME = {
  AGREEMENT: 'agreement',
  FALSE_NEGATIVE: 'false_negative',
  FALSE_POSITIVE: 'false_positive',
  INCONCLUSIVE: 'inconclusive',
  MIRROR: 'mirror',
};

export function classify(expected, actual) {
  if (!VERDICTS.includes(expected) || !VERDICTS.includes(actual)) {
    throw new Error(`Unknown verdict: expected=${expected}, actual=${actual}`);
  }

  if (expected === actual) return OUTCOME.AGREEMENT;
  if (expected === 'escalate') return OUTCOME.MIRROR;
  if (actual === 'escalate') return OUTCOME.INCONCLUSIVE;

  return expected === 'reject' ? OUTCOME.FALSE_NEGATIVE : OUTCOME.FALSE_POSITIVE;
}
                

The order of the checks carries the logic. Equality comes first, which covers all three matching pairs, including the model escalating when escalation was the expected answer. Next comes the mirror case: if the expected verdict is escalate and the model did anything else, it decided where it should have hedged, whichever way it decided. This check has to come before the directional ones. Otherwise a model that rejected a listing that should have gone to a human would be filed as a false positive, which is wrong, since that listing was never compliant or non-compliant to begin with. After that, a model that answered escalate where a decision was expected is inconclusive. Only the two remaining pairs are directional errors. Here the positive class is a flagged violation, so approving something that should have been rejected is a false negative and rejecting something that should have been approved is a false positive.

The guard at the top keeps the five outcomes exhaustive. Anything that is not one of the three verdicts, such as a malformed response, is a bug in the harness and not a decision to be counted, so it fails loudly. It also makes this one of the few parts of an AI test setup that can be tested exhaustively. There are only nine possible pairs: three agreements, two mirror cases, two inconclusive decisions, one false negative, and one false positive. A test that asserts all nine pins the definitions down for good.

Applying it to a whole run is a matter of expanding each test case into its decisions, classifying each, and counting. The case-level result falls out of the same records, since a case is red if at least one of its decisions is not an agreement:

                
export function summarize(cases) {
  const decisions = cases.flatMap((c) =>
    Object.entries(c.expected).map(([rule, expected]) => ({
      caseId: c.id,
      rule,
      outcome: classify(expected, c.actual[rule]),
    })),
  );

  const counts = {};
  for (const { outcome } of decisions) {
    counts[outcome] = (counts[outcome] ?? 0) + 1;
  }

  const redCases = new Set(
    decisions.filter((d) => d.outcome !== OUTCOME.AGREEMENT).map((d) => d.caseId),
  );

  return { total: decisions.length, counts, redCases: redCases.size };
}
                

Each case carries its expected verdicts, like the fixture from earlier, plus the actual verdicts the model returned for the same rules. The red case count is derived from the decisions and is not measured separately, which is what lets the two views be reconciled instead of argued about.

Reconciling the Numbers

At this point the new model's run has been described with two sets of numbers that look like they contradict each other. By cases, 87 of 120 passed, a pass rate of 72.5%, which reads as more than a quarter of the suite failing. By decisions, 442 of 480 agreed with the expected verdict, which reads as under 8% wrong. A reader is right to distrust two figures that tell different stories about the same run, so here is the output of the summarizer for it:

{
  total: 480,
  counts: {
    agreement: 442,
    inconclusive: 30,
    false_positive: 3,
    mirror: 3,
    false_negative: 2
  },
  redCases: 33
}

Both views are in that one object, and they are the same data. The table below walks from one to the other for the new model, with the old model's run beside it. Every decision in a green case is an agreement by definition, so the split that matters is what the red cases contain:

Measure Old model New model
Green cases 101 87
Red cases 19 33
Decisions in green cases (cases × 4) 404 348
Decisions in red cases (cases × 4) 76 132
Agreements inside red cases 55 94
Wrong decisions inside red cases 21 38
Total agreements 459 442
Total decisions 480 480
Pass rate (green cases / 120) 84.2% 72.5%
Agreement rate (agreements / 480) 95.6% 92.1%

Take the new model. The 33 red cases hold 33 × 4 = 132 decisions. Of those, 38 are wrong and the other 94 are agreements, so the average red case is about 71% right. The 87 green cases hold 87 × 4 = 348 decisions, all of them agreements. The agreement count therefore adds back up as 348 + 94 = 442, and the full set adds up as 442 + 38 = 480. The old model does the same: 404 + 55 = 459 agreements, and 459 + 21 = 480 decisions. The two figures disagree only because they are dividing by different things. A case is red when any one of its four decisions is wrong, so a few wrong decisions are spread across many red cases while each red case still contains mostly correct ones.

There is one more thing to square. The new run has 33 red cases but 38 wrong decisions, so five of the wrong decisions are not accounted for by "one per red case". They sit in cases that hold more than one miss: 29 red cases contain a single wrong decision, 3 contain two, and 1 contains three, which is listing 071 from earlier. Then 29 + 3 + 1 = 33 cases and 29 + (3 × 2) + (1 × 3) = 38 decisions. The old run has 17 red cases with one miss and 2 with two, which gives 19 cases and 17 + 4 = 21 decisions. This is the gap between the two counts, written out.

The same 38 wrong decisions then split by outcome as 2 false negatives, 3 false positives, 30 inconclusive, and 3 mirror cases, and 2 + 3 + 30 + 3 = 38. A run can be read as cases, decisions, or outcomes, and each level adds up to the one above it. Nothing was added or dropped along the way.

Note what this reconciliation does not do. The agreement rate fell from 95.6% to 92.1% too, so counting decisions instead of cases does not by itself rescue the new model. Whether more mismatches are worse depends on which decisions they were, and that is the question the outcome breakdown and its prices exist to answer.

Why a Safer System and a Worse One Look the Same

Here are the wrong decisions of both runs, split by outcome and priced with the figures from the earlier table:

Outcome Price each Old count New count Old cost New cost
False negative $400 9 2 $3,600 $800
False positive $15 4 3 $60 $45
Inconclusive $4 8 30 $32 $120
Mirror case $15 0 3 $0 $45
All wrong decisions – 21 38 $3,692 $1,010

Read by outcome, the new model is a clear improvement. Seven fewer violations were waved through, one fewer compliant listing was rejected, and the cost of the misses fell from $3,692 to $1,010, a drop of about 73%. The price was 22 more inconclusive decisions and 3 mirror cases that did not exist before. The wrong decisions went up by 17 in total, and that is exactly −7 − 1 + 22 + 3. More mistakes, at a much lower price.

Now consider a hypothetical third model that really is worse. It makes the same 17 extra mistakes, but all of them are violations waved through, so its false negatives rise from 9 to 26 while the rest stays as it was. If those misses fall across the suite in the same pattern, it also has 33 red cases and 442 agreements. Next to the other two runs it looks like this:

Model Pass rate Agreements Cost of misses
Old model 84.2% 459 $3,692
New model 72.5% 442 $1,010
Hypothetical worse model 72.5% 442 $10,492

The last two rows have the same pass rate and the same agreement count. One of them costs about a tenth as much as the other. Nothing on the dashboard separates them, and a team that sees only the red count would react to both in the same way.

Collapsed into one red count, a system that got safer is indistinguishable from one that got worse.

The damage is not limited to a wrongly blocked release. A pass rate that is treated as the goal pulls the team in a particular direction. The cheapest way to bring the number back up is to make the model more decisive, which turns hedges into confident answers. Some of those answers will be right and the count will improve. The ones that are wrong will land in the expensive direction, because a model that stops escalating stops catching its own doubtful cases. The metric rewards the change that removes the safety margin.

What to Report Instead of a Pass Rate

None of this calls for a complicated scoring system. It calls for reporting a run as the outcomes it contains, and for gating on the outcomes that cost the most:

The pass rate can stay as a diagnostic. The list of red cases is still a good place to start looking at what changed. It just should not be the number that decides a release.

Takeaways