h.work
Operations

What oversight actually looks like at ten deployments

One expert can attend AI Specialists across ten to thirty companies — but only if oversight is designed in advance.

Shaky Spears · Sep 1, 2026 · 4 min read
What oversight actually looks like at ten deployments

What oversight actually looks like at ten deployments

Most companies define oversight of an AI system in roughly the same way: someone checks the work. A person sits between the system's output and its consequences, reviews what came out, and either approves it or sends it back.

That answer holds for one deployment. It stops holding at ten.

A credentialed expert in the h.work consortium can attend AI Specialists working for somewhere between ten and thirty companies at once. That number is not a productivity claim; it is a design constraint. If oversight meant reading everything, the number would be one. The only way a single senior practitioner covers thirty companies is if most of the work never reaches them at all — and if the decisions that do reach them were identified before anything went wrong.

So oversight at scale is not a volume of reading. It is a set of decisions made in advance about where human judgement is load-bearing.

The two questions that define an oversight practice

Every deployment resolves into two questions, and the answers are rarely obvious.

Which decisions are consequential? Not important — consequential. A late reply to a customer is important. A credit note issued against the wrong account is consequential, because it moves money, touches the ledger, and is expensive to unwind. The distinction matters because most operational work is important and almost none of it is consequential. Teams that conflate the two either escalate everything, which collapses the model, or escalate nothing, which is what people mean when they say autonomous AI feels risky in high-stakes work.

What does the expert need in order to decide? A reviewer who receives a flagged item with no context has to reconstruct the case before ruling on it, and reconstruction is where the time goes. An expert holding thirty companies cannot afford reconstruction. The escalation has to arrive with the decision already framed: what was found, what the rule says, which two options exist, what happens under each.

Answer both well and one person genuinely supervises thirty deployments. Answer either badly and the oversight layer becomes the bottleneck the AI was meant to remove.

What the expert's day looks like

It is quieter than people expect, and less reactive.

Most of the volume passes under continuous monitoring rather than review. The Specialist works inside the company's existing channels — Slack, Teams, email, the ERP — and its actions are logged. Monitoring watches the shape of that log for drift: a category of exception appearing more often than it used to, a step being skipped, a tolerance being approached repeatedly rather than breached once.

A much smaller stream arrives as decisions. Those are the pre-identified consequential calls, routed to the expert before execution, not after. This ordering is the whole argument. Review after execution is auditing. Review before execution is oversight, and only the second one prevents anything.

The third stream is the one that does not appear in most vendor diagrams: correction. When the expert overrules a Specialist, the correction is not filed as an exception and forgotten. It becomes ground truth. The case that required a human this month is the case that should not require one next quarter, because the boundary has been redrawn with the expert's reasoning inside it.

That is the part that compounds. The first two streams keep a deployment safe. The third is what makes the tenth deployment cheaper to attend than the first.

Why this is a staffing question, not a tooling one

Mid-market operators recognise the underlying problem long before they recognise the model. Senior judgement in their business is scarce, expensive, and concentrated in a small number of people. Those people spend most of their week on work that does not require their seniority — reconciling, chasing, checking, formatting — and the fraction that genuinely needs them gets whatever attention is left.

Automating the throughput without addressing the judgement does not fix this. It produces more output arriving at the same narrow point, faster. The constraint was never the volume of work the company could produce. It was the volume of consequence a small number of qualified people could responsibly hold.

Which is why the interesting question about an AI workforce is not how capable the Specialists are. It is how the attention around them is arranged: who is accountable, what reaches them, in what form, and whether that person is verifiably who and what they claim to be.

Verification is doing real work here

Named oversight only means something if the name resolves to a real, credentialed person. An oversight layer of anonymous reviewers is a promise, not a control. This is where Humanity sits underneath h.work rather than beside it: the consortium's experts are identity-verified and their credentials — CPA, JD, MD, CAMS, and others — are cryptographically attested through Humanity's verification infrastructure, provable without exposing the underlying personal data.

The practical consequence for a company is narrow and specific. You can know who supervises your Specialists. You can know what they are qualified in. You can see the audit trail of what they reviewed and what they changed. "Powered by humanity" is describing that mechanism, not decorating it.

What to ask a vendor

If you are evaluating any supervised-AI arrangement, three questions separate a designed oversight practice from a claim.

First: name the decisions that route to a human, before deployment. A vendor who cannot list them for your workflow has not done the design work; they will do it after your first incident.

Second: ask what the reviewer sees. If the answer is a queue of flagged items, you are buying an auditing function. If the answer is a framed decision with options and consequences, you are buying oversight.

Third: ask what happens to a correction. If it closes a ticket, quality is flat over time. If it becomes training signal, the deployment gets quieter as it ages — and the expert's capacity to attend the next ten companies goes up rather than down.

Attention is the scarce input. Everything about how it is spent is a design decision, and almost all of those decisions get made before the first task runs.