A paper-craft illustration of a whale and a school of fish moving together through a coral reef, lit by rays of light from the surface above, representing a mostly self-sufficient system still being watched from above
AI & Trust

Human Intervention Rate: The One Metric That Tells You If Your AI Rollout Is Actually Working

Most companies running AI agents in production cannot tell you how often those agents actually need a person to step in. Human Intervention Rate is the number that answers that question — and by the end of this piece, you should be able to calculate it for a workflow you already run.

FabricLoop Editorial
2,180 words
10 min read

FabricLoop's own framework for AI organizations defines Human Intervention Rate plainly: it asks how often automated work needs a person. That is the definition, and this piece does not deviate from it. What follows is the part the concept page does not spell out in full: the actual arithmetic, applied to one real workflow, with the numbers that make the idea concrete instead of aspirational.

What the number actually measures

Human Intervention Rate (HIR) is the share of an agent's actions, within one defined workflow and one time period, that required a person to step in before the outcome could stand as finished. "Step in" has a specific meaning here: a person corrected the output, overrode a decision the agent made, or answered a question the agent explicitly raised before it would proceed — what FabricLoop's Loop Agent calls an ask_human moment. Divide the count of those actions by the total number of actions the agent took in the same period, and you have HIR.

The reason this metric earns a place next to uptime and accuracy, rather than underneath them, is that it measures something those numbers cannot see. An agent can post a 95% accuracy score on some internal benchmark and still be a worse rollout than one scoring 80%, if the 5% it gets wrong slips through silently while the 20% it's unsure about gets flagged every time. HIR does not ask whether the agent is good. It asks whether the system knows when it needs a person, and whether a person actually shows up when it does. That second question is the one that determines whether a rollout is safe to expand.

Calculating HIR for one real workflow

Take a workflow an IT or operations team might actually run today: an agent that triages incoming support tickets, classifies them (billing, bug report, refund, account access, and so on), and drafts a first-pass reply. Every draft lands in a review queue before it reaches a customer — nothing sends on its own. That review step, by itself, is not intervention. A reviewer clicking "send" on a draft that needed no changes is the workflow working as designed. Intervention is what happens when the draft needed work: a reviewer rewrote it, corrected the classification, redirected the ticket to a different queue, or the agent itself paused mid-task and asked a question before drafting anything.

The numbers below are an illustrative example, not a real company's data — but the shape of the story, and the arithmetic behind it, is exactly what you'd build from your own logs.

HIR Formula
HIR = Actions Requiring Intervention ÷ Total Agent Actions
Same workflow, same period. Count only actions where a person changed the outcome. A reviewer approving an unread draft doesn't count as intervention — and neither does a draft that needed no changes at all.
Escalation Share
Escalation Share = Agent-Initiated Asks ÷ Total Interventions
Splits every intervention into two kinds: the agent flagged its own uncertainty, or a reviewer caught a mistake the agent didn't flag. This is the number that tells you whether a falling HIR is good news.

In the pilot month, the agent touches 640 tickets. Of those, 415 need an intervention — a rewrite, a reclassification, or a reroute — and only 75 of those 415 are moments the agent flagged itself before drafting anything. The rest are mistakes a reviewer catches after the fact. That's an HIR of 64.8%, with an escalation share of just 18%: the agent is confidently wrong most of the time it's wrong, which is the worst version of this problem to have.

The team pulls the correction log and tags each intervention with a reason. Two categories dominate: the agent misreads refund policy on anything involving a dollar amount, and it drafts calm, procedural replies to customers who are visibly angry. Both are fixable without touching the model — add an explicit rule that any ticket mentioning a refund over $50, or scoring above a sentiment threshold, triggers an ask_human escalation instead of a draft. Everything else still gets drafted and reviewed as before.

MonthTickets HandledInterventionsHIREscalation Share
1 — Pilot 640 415 64.8% 18%
2 — After rules added 810 224 27.7% 58%
3 — Rules tuned again 940 101 10.7% 79%

By month three, HIR has dropped by more than 80%, but the more informative number is the escalation share: it climbed from 18% to 79%. Most of what's left isn't the agent getting caught being wrong — it's the agent correctly recognizing a genuinely ambiguous case (a VIP account, a policy exception, a refund that falls right at the threshold) and asking before it acts. The decline is real, and it's earned: each round of corrections got fed back into explicit rules, so the specific mistakes that produced them stopped recurring, while the categories that still need judgment keep getting flagged instead of drafted around.

The decline that matters is the one where the agent gets better at knowing what it doesn't know — not the one where a person quietly stops checking.

The mistake: treating zero as the goal

Once a team is watching HIR fall month over month, the obvious next question is how low it can go. The instinct is to treat zero as the finish line — proof the agent finally got good enough to run unsupervised. That instinct is backwards, and it's the most common misreading of this metric.

Why 0% is usually a red flag

A workflow showing 0% intervention for weeks on end almost never means the agent stopped making mistakes. It means one of two things happened instead: reviewers stopped actually reading the drafts before approving them, or the escalation path quietly broke — thresholds got loosened, a routing rule failed silently, or the ask_human trigger stopped firing. Either way, the zero isn't telling you the system stopped needing a person. It's telling you a person stopped being asked, or stopped looking.

The actual goal was never fewer interventions in the abstract. It's a system where the specific moments that need a person's judgment get surfaced — and only those moments — so that a person's attention goes to what actually needs it, instead of being split evenly across everything or missing entirely. A workflow sitting at 12% HIR, where nearly all of that 12% is the agent correctly flagging genuinely ambiguous or high-stakes cases, is healthier than one sitting at 2%, where most of that 2% is a reviewer stumbling onto an error the agent never flagged. The lower number can hide the worse system.

This is exactly what escalation share is for. Watched next to HIR, it tells you which story you're in:

Reading an HIR trend — the same falling number, two different meanings
0%, indefinitely
Red flag — nobody's watching, not a flawless system
High and flat for months
Not learning — corrections aren't feeding back into rules
Falling, escalation share falling too
Check it — likely rubber-stamped approvals, not real progress
Falling, escalation share rising
Trust earned — the system knows its own edges

If HIR is dropping while escalation share holds flat or falls, don't file it as a win yet. Pull a random sample of the actions logged as "no intervention needed" and have someone review them cold, without telling them the sample was flagged clean. Check whether downstream signals — tickets that get reopened, complaints, refund clawbacks, CSAT — are drifting upward at the same time. A falling HIR with rising downstream problems is not a system that learned faster. It's a system nobody caught in time.

What to instrument if you want to measure this today

None of this requires new tooling so much as it requires logging the right thing. Most teams running an agent already track volume — how many tickets it touched, how many tasks it drafted. Almost none of them track outcome, which is the only thing HIR actually needs.

  1. Log an outcome for every action, not just an activity count. Sent as-is, edited before sending, rejected and rewritten, or escalated by the agent itself. Without outcome-level logging, HIR can't be computed at all — you'll know the agent did something, not whether it needed to be fixed.
  2. Fix your denominator before you fix your numerator. Decide what counts as one action for this workflow — one ticket touched, one task drafted — and hold that definition steady across periods, so a change in HIR reflects the agent's judgment and not a change in how you're counting.
  3. Tag every intervention with a reason. "Edited" tells you almost nothing. "Edited: misapplied refund policy above $50" tells you exactly what to fix next. A short, consistent taxonomy turns a correction log into a punch list instead of a scoreboard.
  4. Track escalation share alongside HIR, not instead of it. The two numbers together tell you whether a decline is earned or borrowed — see the trend table above.
  5. Set a floor, not a target of zero. Decide, per workflow, what a plausible non-zero HIR looks like given how much real ambiguity that workflow contains, and treat a rate that drops well below that floor as something to investigate, not something to celebrate.
  6. Report HIR per workflow, never as one blended company-wide number. A single average hides which specific workflow has actually earned less oversight and which one is quietly accumulating risk underneath a good-looking headline figure.
  7. Re-check the "clean" sample on a schedule. Periodically pull actions logged as needing no intervention and have someone review them without knowing they were flagged clean. It's the only direct check on whether your reviewers are still reading.
FL
How FabricLoop supports this

This is why Loop Agent is built around ask_human, resume, and channel-app escalation rather than silent autonomy — an agent that pauses to ask is an agent that shows up in your HIR numerator on purpose, not one that got caught by accident. Escalations and drafts surface in the same Groups where the team already works, next to tasks and notes, so the moment that needed a person is visible where the work already lives — not buried in a separate agent console nobody checks. On Enterprise, audit logs let IT and ops see what agents did and exactly when a human stepped in, which is the raw material HIR is built from in the first place.

Pair this with Legibility — the companion concept for making sure grants and access are visible too — and you get the two questions every AI rollout should be able to answer before it expands: who can see what an agent is doing, and how often does a person actually need to step in.


Key takeaways
01
Human Intervention Rate is the share of an agent's actions, in one workflow and one period, that required a person to correct, override, or answer a question the agent raised before the work counted as done. It's a ratio: interventions divided by total actions.
02
A reviewer approving a draft that needed no changes is not an intervention. HIR measures how often the outcome had to change, not how often a human looked at something.
03
Escalation share — the portion of interventions the agent flagged itself, versus the portion a reviewer caught after the fact — is the companion metric that tells you whether a falling HIR reflects real improvement or fewer people actually checking.
04
A healthy decline in HIR comes from feeding correction reasons back into explicit rules or examples, so the same mistake stops recurring — not from reviewers getting tired of reading drafts.
05
0% intervention sustained over time is almost always a red flag, not a milestone. It usually means reviewers stopped reading or an escalation path broke silently — not that the agent became flawless.
06
The actual goal is not the lowest possible number. It's a system that makes the specific moments needing human judgment visible — and only those moments — so a person's attention lands on what actually needs it.
07
To sanity-check a falling HIR, watch downstream signals — reopened tickets, complaints, refund clawbacks, CSAT — for a rise that the rate alone would hide, and periodically re-review a sample of "no intervention needed" actions cold.
08
To measure HIR at all, you need outcome-level logging (sent as-is, edited, rejected, escalated) — not just activity counts. Most teams running agents today log volume and nothing else.
09
Report HIR per workflow, not as one blended company-wide figure. A single average can hide a workflow that's actually earned less oversight next to one that's quietly accumulating risk.