h.work
Operations

The Escalation Rule Nobody Wrote Down

Most ops teams run automated work without a written rule for what a human must approve.

Shaky Spears · Aug 27, 2026 · 4 min read
The Escalation Rule Nobody Wrote Down

The escalation rule nobody wrote down

Ask an operations lead which decisions in their workflow require a human sign-off and you will usually get a confident answer. Ask them where that rule is written and the answer changes shape. It is in someone's head. It is in a Slack thread from March. It is in the way the team has always done it.

That gap is not a documentation problem. It is a control problem, and it gets expensive in a very specific way: the moment work moves faster than the people supervising it, an undocumented threshold stops being a judgement call and becomes a coin flip.

What the missing rule actually costs

Automated and AI-assisted work in a mid-market company rarely fails on the routine cases. Routine is what it is good at. It fails at the edges — the refund that is technically within policy but commercially wrong, the supplier substitution that clears spec but breaks a customer's own certification, the invoice that reconciles perfectly and should still never have been paid.

Each of those is an escalation. And in most teams, whether the escalation happens depends on who happened to be looking. Two consequences follow.

The first is inconsistency. The same class of decision gets approved on Tuesday and rejected on Thursday, and nobody can reconstruct why. Customers notice this before management does.

The second is that accountability quietly evaporates. When the routing rule lives in habit rather than in writing, nobody is on the hook for a decision that was never formally routed to them. The work happened. The judgement did not. And when someone eventually asks who approved it, the honest answer is that the question was never asked.

Why "a human reviews everything" is not the fix

The instinctive correction is to put a person in front of every output. It fails for two reasons.

It does not scale, obviously. But more importantly, it degrades the review itself. A reviewer signing off two hundred routine items a day is not exercising judgement; they are clicking. By the time the item that genuinely needed attention arrives, the reviewing muscle has been trained to wave things through. Blanket review produces the appearance of oversight and the reality of none.

The useful question is not how much gets reviewed. It is which things, decided in advance, on criteria you can state out loud.

Write the threshold down

An escalation rule that works has four parts, and all four fit on one page.

1. The trigger. A specific, checkable condition — not "anything unusual." Value above a stated amount. A first-time counterparty. A deviation from spec. An exception to a documented policy, regardless of size. A decision touching a regulated obligation. If a rule cannot be evaluated without a judgement call, it is not a trigger; it is a wish.

2. The reviewer. Not "someone senior." A named role with the standing to say no, and the credential to make that no defensible. This is where most escalation rules fail in practice: the trigger fires, the item lands with whoever is available, and the review is nominal because the reviewer has no authority over the outcome.

3. The window. How long the reviewer has before the decision either proceeds or halts, and which of those two is the default. An escalation rule without a default is a bottleneck waiting for a busy week. State it: this class of item waits, that class proceeds and gets reviewed after the fact.

4. The record. What was escalated, who decided, on what basis, and when. Not for its own sake — because the record is what lets you tune the thresholds later. Without it, you are guessing at whether your trigger is set too tight or too loose forever.

Tune it with the evidence you generate

A first draft of an escalation rule is always wrong. That is fine, as long as it is wrong in writing.

Two numbers tell you what to change. If reviewers are approving almost everything that reaches them without modification, the trigger is too broad — you are spending senior attention on work that did not need it, and training reviewers to skim. If problems are surfacing in items that never triggered a review, the trigger is too narrow, and the pattern in those misses tells you which condition to add.

Reviewing the log monthly and adjusting one threshold at a time gets a team to a defensible rule within a quarter. Nobody gets there by reasoning about it in the abstract.

The judgement layer is a design choice

There is a version of automation where you deploy capable systems, hope the edges are rare, and discover the escalation rule only after an edge case becomes an incident. It is the common version. It is also entirely avoidable.

The alternative is to treat the boundary between throughput and judgement as something you design deliberately. Machines take the volume: the repeatable work, the optimisation inside stated rules, the twenty-four-hour execution. People take the calls where rules do not neatly apply — taste, regulatory nuance, accountability, the cases nobody anticipated. The escalation rule is simply where you draw that line, written down so it holds when the person who invented it is on leave.

This is why h.work pairs every AI Specialist with credentialed human oversight rather than shipping capability alone. Consequential decisions route to a senior expert whose credentials are verified through Humanity before execution; routine work runs continuously and their corrections feed back into how the Specialist handles the next case. The threshold is explicit, the reviewer is named and qualified, and the record exists — because oversight that is not specified is not oversight.

Start here

You do not need a project to fix this. Take the workflow that would cause the most damage if it went wrong unsupervised, and answer four questions in writing: what triggers a human review, who performs it, how long they have, and where the decision gets recorded.

If you cannot answer all four today, you already know which one to write first.