TL;DR
- The core decision: what the reviewer of an AI output is actually there to contribute.
- When doing nothing is right: when the reviewer holds something the feature can't and uses it.
- What has to be true: they can disagree, and something would happen if they did.
- How the options split: by what the reviewer is shown and whether they can overturn anything.
- Decision rule: give them what the model lacked, or you've added a signature and no check.
- Outcome to expect: fewer reviews, each capable of producing a different answer.
There's a Human Reviewing Everything
You ask what happens if the AI gets something wrong, and you're told a person reviews every output before it goes anywhere. It's a reassuring answer. It's also the answer everybody gives, and it's almost never examined.
So examine it. Somebody is handed an output. It reads well, it's formatted like every other output, and it contains nothing indicating whether it had any basis. They have a queue of them and a day's other work. What, specifically, are they meant to do?
If the honest answer is read it and see whether it looks reasonable, then the review catches only outputs that look unreasonable, and the errors that matter in this area don't. A confident wrong answer reads exactly like a confident right one. Asking somebody to spot the difference by inspection is asking them to do something the task doesn't permit.
Best tools for AI HR Tools
That's the reframe. The reviewer isn't there to be careful. Care doesn't help, and treating the problem as insufficient diligence produces reviewers who feel bad and outputs that are just as wrong. The reviewer is there because they hold something the feature doesn't: knowledge of this person, of your organisation, of what changed last month, of what the question was really asking. A review that draws on that is a genuine control. A review that doesn't is a signature.
Everything below follows from that. What the reviewer knows, what they need to be shown to use it, and what has to be true for their disagreement to change anything.
One boundary. Whether human oversight is required in your situation, what form it must take, and what you owe the person affected are questions whose answers differ by jurisdiction and are changing quickly. Nothing here tells you what applies to you. Establish it with local advice, and take the governance framing from the AI in the workplace material.
When You Genuinely Do Not Need to Act Yet
The reviewer holds relevant knowledge and uses it. They know the people or the subject, they're shown enough to apply it, and they sometimes say no. That's a working arrangement.
Review happens but nobody ever disagrees. Worth examining rather than celebrating. A review that has never changed an outcome is either reviewing something that's never wrong or isn't reviewing.
The reviewer can't judge what they're given. They see the output, they have no basis for an opinion, and they approve. That's the common case and it's the one to fix.
The edge case that forces it. Something wrong went out and the review record shows it was approved. The record then makes things worse rather than better, because it documents a control that wasn't operating.
Five Questions This Reader Asks at 11pm
What is the reviewer actually checking? The question to answer concretely, per feature, in a sentence. Not reviewing the output, which means nothing. Checking that the named specifics appear in the source, or that this matches what they know of the person, or that the question was understood. Something a person can do and know when they've done it.
Can they tell a good output from a bad one? Frequently not, on inspection alone, and this isn't a failing on their part. If the feature produces the same fluent formatting whatever the basis, then judging by appearance is impossible and the review needs to draw on something else.
Is checking a sample enough? Depends entirely on why you're checking. To find out whether a feature is broadly reliable, a sample is fine and a good use of time. To prevent a wrong output reaching a particular person, it does nothing, because the one that reaches them is the one you didn't sample.
Does review make AI safe to use? Not by existing. It helps to the exact extent the reviewer holds something the feature lacked and is positioned to act on it. Review as a formality provides no protection and does something worse: it produces a record suggesting protection, which stops anybody asking.
What if the reviewer would rather not approve? Then it matters enormously what their options are. If the only alternative to approving is stopping the whole process, they'll approve. A way to query one item, escalate it, or hold it while everything else proceeds is what makes refusal a real possibility.
What the Reviewer Has That the Model Does Not
| The knowledge | Why the feature can't hold it | What it catches |
|---|---|---|
| This particular person's situation | It isn't in any record it reads | General statements misapplied to an exception |
| What changed recently | Its sources haven't caught up | Confident answers from stale information |
| What the question really meant | It has only the words it was given | Correct answers to a different question |
| What your organisation actually does | Practice diverges from documentation | Answers true on paper and wrong in reality |
| What isn't written down anywhere | It reads records, and this isn't in them | The arrangement everybody knows and nobody recorded |
| How this will land with the reader | It has no sense of the relationship | Technically right and badly received |
| What was decided but not yet recorded | The record hasn't been updated | Answers overtaken by a decision already made |
| Whether the answer is plausible here | It has no baseline for your organisation | Figures that would be normal elsewhere |
The first row is the whole argument for human review in HR specifically. Almost every general statement about employment has exceptions attached to particular individuals, those exceptions live in contracts, side agreements and institutional memory, and a feature reading standard records has no way to know they exist or that it should hesitate.
The fifth row is the one people underestimate. A surprising amount of how any organisation works isn't recorded anywhere: the arrangement with that team, the thing that was agreed verbally, the practice that superseded the policy. A reviewer who has been there a while carries this, and it's invisible to anything reading records.
The last row is the cheapest check available and the most transferable. Somebody familiar with your organisation knows what a normal figure looks like, and an output producing something that would be unremarkable somewhere else and is impossible here gets caught by a person who has simply seen a lot of these.
Five Diagnostic Questions You Can Self-Assess Against
What, in one sentence, is the reviewer checking? If you can't write it, neither can they, and what's actually happening is a glance and an approval.
Are they shown what the output was based on? Judging an output without its source is guessing at whether it's well founded. Showing the source is usually a configuration option nobody switched on.
Has a review ever changed an outcome? If never, that's the finding. Either nothing has ever been wrong, which is unlikely, or the review isn't functioning.
What happens if they say no? Trace it. If the answer is that the process stops for everybody, nobody will ever say no, and that's a design problem rather than a courage problem.
Does the reviewer know the people or the subject? If outputs about individuals are reviewed by somebody who doesn't know them, the main thing a human could contribute isn't available.
Ask all five of the reviewer rather than the process owner. The two descriptions of the same review routinely differ, and the reviewer's version is the one that describes what happens.
Six Review Arrangements, Reviewed
A reviewer who sees the output alone
They're shown what the feature produced, and nothing else. It earns its place as better than nothing: obvious nonsense gets caught, and formatting problems get caught.
Where it falls short is everything else. Without the source, without the question as asked, without the record it drew on, the reviewer is assessing plausibility, and plausibility is precisely what these features are best at. Confident wrongness passes every time.
The most common arrangement and the weakest. Showing the source alongside the output is frequently a setting rather than a project.
Where the source genuinely can't be shown, the honest response is to narrow what the feature is used for rather than to keep the review and call it a control. A reviewer who can't check anything is better redeployed than left approving things.
It's also worth asking the vendor directly whether the underlying material can be surfaced. Plenty of products can and don't by default, because the default was chosen for a clean interface rather than for reviewability.
A reviewer who sees the output alongside what it was based on
They get the output and the material behind it. It earns its place as the arrangement that makes review meaningful, because now the reviewer can check rather than assess, and checking is a different activity with a different success rate.
Where it falls short is time. Comparing against a source takes longer than reading an output, so it needs resourcing honestly, and under pressure it silently degrades into the previous arrangement.
The right default. If the volume makes it impossible, that's information about the volume rather than a reason to drop the source.
Watch for the slow degradation, because it's invisible from outside. Nobody announces that they've stopped opening the source; they just start with the ones that look unusual, then only the ones that look unusual, and within a couple of months the arrangement has quietly become the previous one.
A short note recorded per review, saying what was checked, keeps this honest without much effort. It also gives you something to look at later when you're asking whether the control was operating.
A reviewer who only sees the cases the system was unsure about
The feature flags what it couldn't handle confidently, and a person handles those. It earns its place by concentrating human attention where it's most likely to be useful, which is a genuine improvement over reviewing everything shallowly.
Where it falls short is that the feature's uncertainty and its wrongness aren't the same set. The dangerous output is the one it was confident about and wrong, which by definition never reaches the reviewer. The flagged set is useful, and it isn't the risky set.
Good alongside something else. Treat it as triage rather than as the whole control, and sample the unflagged cases too.
The reasoning is worth being explicit about with whoever designed the arrangement, because flagging feels like it solves the problem and it addresses a different one. Uncertainty flags tell you where the feature knew it was struggling. The output that causes harm is the one where it didn't know.
Ask as well what the flagging threshold is and who set it. If it can be tuned, somebody has tuned it, and the setting determines how much reaches a person at all.
A reviewer who checks a sample after the fact
Somebody examines a proportion of outputs after they've gone out. It earns its place as the only practical way to learn about a feature's general behaviour at volume, and it's how you find patterns.
Where it falls short is prevention, which it doesn't do at all. Anything wrong in the unsampled majority has already reached its destination, and anything found in the sample is found after the fact. It's a measurement instrument rather than a control.
Valuable for what it's for. Don't let it stand in for preventing a wrong output reaching somebody, because those are different jobs.
The sample also needs to be drawn properly. Reviewers who choose which outputs to examine will tend, reasonably, towards the interesting ones, and a sample selected that way describes the interesting cases rather than the ordinary ones where quiet errors live.
Record what the sample found, each time. The value of after-the-fact checking accumulates across rounds, and it accumulates only if somebody wrote down what turned up.
A reviewer with no authority to overturn the result
They look, they can flag, and they can't stop anything. It earns its place only as an early-warning mechanism, and only if somebody acts on what they flag.
Where it falls short is obvious once stated and common in practice. A reviewer who can't change an outcome isn't a control, they're an observer, and their approval in a record misrepresents what happened. It's also corrosive: people asked to review things they can't affect stop reviewing them properly, reasonably.
If review matters, give the reviewer authority. If it doesn't, don't record their approval as though it were a check.
The middle position that sometimes works is a reviewer who can hold an item pending somebody else's decision. They aren't overturning anything themselves, and they can stop something long enough for a person who can to look at it, which is meaningfully different from flagging into a queue.
What makes that work is a named person at the other end with a stated response time. Without one, holding an item is indistinguishable from losing it, and reviewers learn quickly not to.
Review that exists in the process description and not in anybody's week
The documentation says outputs are reviewed. Nobody's actual working day contains it. It earns its place nowhere and it's more common than anybody would like.
Where it does damage is the gap between the described process and the real one, which nobody notices until an incident, at which point the documentation says something happened that didn't.
Worth checking rather than assuming. Ask the named reviewer how long they spend on it, and compare that with the volume.
The discrepancy usually isn't anybody's fault. A process written when volumes were lower carries on being the documented process long after the volumes changed, and nobody revisits it because nothing has visibly broken.
The risk is what the document does during an incident. A process description stating that every output is reviewed becomes the standard you'll be measured against, including by yourself, and discovering it was aspirational is a bad moment to have in public.
The Decision Table
| Situation | Scale | Setup | Primary Pain | Recommended Starting Point |
|---|---|---|---|---|
| Reviewer knows the subject and disagrees sometimes | Any | Any | None | Change nothing |
| Reviewer sees the output only | Any | Any | Assessing plausibility | Show them the source |
| Reviewer doesn't know the people | Any | Centralised | The main contribution unavailable | Move review to somebody who does |
| Review has never changed anything | Any | Any | A signature, not a check | Find out which of the two it is |
| Only option is approve or stop everything | Any | Any | Refusal isn't realistic | A path to query one item |
| Reviewer can flag but not overturn | Any | Any | An observer recorded as a control | Give authority, or stop recording approval |
| Only flagged cases reviewed | Any | Any | Confident errors never surface | Sample the unflagged too |
| Sampling used to prevent errors | Any | Any | Wrong job for the instrument | Add a check before the output lands |
| Review in the document, not the week | Any | Any | Described process isn't the real one | Ask the reviewer what they do |
The fourth row is the diagnostic that settles everything else. A review that has never produced a different outcome is either unnecessary or not happening, and finding out which takes one conversation with the person doing it.
The fifth row is the design flaw that produces rubber stamping, and it gets misdiagnosed as a people problem every time. Somebody whose only alternative to approving is halting a process that everybody depends on will approve, indefinitely, and no amount of asking them to be thorough changes that arithmetic.
The third row is the one that gets introduced deliberately, for good reasons, and quietly removes the control. Centralising review looks like consistency and specialisation, and it moves the work away from the only people who could recognise an exception.
Review That Is a Signature
The reviewer needs a basis, not an instruction. Telling somebody to check an output carefully gives them nothing to check against. Telling them to confirm that the named figures appear in the attached source gives them a task they can complete and know they've completed.
Volume determines depth, whatever the policy says. A reviewer with more outputs than minutes will develop a faster method, and that method will be reading for plausibility. This happens within weeks and it isn't a discipline failure, it's arithmetic.
Predictability kills attention. A review that arrives at the same time, in the same form, and has been fine every time before will eventually be approved without being read. Anything that varies, or that occasionally genuinely needs intervention, holds attention far better.
A long run of correct outputs is not reassurance. It's the condition under which attention decays fastest, because nothing in the reviewer's experience suggests the next one needs looking at.
No realistic alternative means no real review. If saying no is expensive, disruptive or unsupported, approval is the only reasonable behaviour available. The fix is a path to hold or query one item without stopping everything.
A record of approval is not evidence of review. This is the part that makes a nominal review worse than none: it produces documentation of a control, which prevents anybody asking whether the control exists. The record outlives everybody's memory of what actually happened.
Recording what was checked fixes most of this. A line naming the specific thing the reviewer verified is more useful afterwards than a tick, and it's considerably harder to produce without having done it.
The right reviewer beats a better process. Somebody who knows the person, the subject or the history will catch things no process design can, and somebody without that knowledge won't catch them however well the process is arranged. Placement matters more than procedure.
Continuity is part of the qualification. Much of what makes a reviewer valuable is accumulated rather than trained: what this team is like, what was agreed years ago, which figures have always looked odd. That takes time to build and leaves with the person, which is worth knowing before review is treated as an interchangeable task.
The uncomfortable conclusion is that most human review in this area is doing less than the people relying on it believe. That's fixable and it's mostly fixable by changing what the reviewer is shown and what options they have, rather than by asking anybody to try harder.
Both of those are configuration and process decisions rather than cultural ones, which is the encouraging part. They can be changed this quarter by whoever owns the system, without anybody's behaviour needing to improve.
Where These Arrangements Go Wrong
| The failure | How it shows up | What would have to change |
|---|---|---|
| Output shown without its source | Plausibility assessed, not checked | Show what it was based on |
| Review centralised away from knowledge | Exceptions invisible to the reviewer | Put it in front of somebody who knows |
| No realistic way to refuse | Approval every time, indefinitely | A path to query one item |
| Flagged cases only | The confident errors never arrive | Sample the unflagged |
| Sample check used as prevention | Errors found after they landed | A check before it reaches anybody |
| Approval recorded without review | Documentation of a control that isn't | Either make it real or stop recording it |
The first row is the highest-value fix on this list because it's usually a setting rather than a project. Whether the reviewer sees the underlying material alongside the output changes the activity from judging to checking, and those have very different success rates.
The sixth row is the one with consequences beyond the immediate error. An approval record is produced, filed, and relied upon by people who weren't there, and it says a person checked this. When that turns out not to describe what happened, the problem is no longer a wrong output.
The third row deserves reading alongside it, because the two combine badly. A reviewer with no realistic way to refuse produces exactly the record described above, every time, and the record is what everybody will look at afterwards.
What to Put in Writing
| Artefact | Who owns it | When it is written | What it prevents |
|---|---|---|---|
| What the reviewer checks, in one sentence | Whoever owns the process | Before relying on review | An instruction nobody can act on |
| What they're shown alongside the output | Whoever owns the system | Before rollout | Plausibility standing in for checking |
| Who reviews, and what they know | Whoever owns the process | Before rollout | Review by somebody without context |
| What happens when they say no | Whoever owns the process | Before rollout | Refusal that isn't realistic |
| Expected volume per reviewer per day | Whoever owns the process | Before rollout | Depth quietly degrading |
| What applies to you, per jurisdiction | You, with local advice | Before decisions about people | A position you assumed |
The fifth row is the one that predicts whether any of the others survive. A review designed for a volume nobody checked against a real working day will be abbreviated within a month, and the process document will continue describing the original version indefinitely.
The fourth row is worth writing because it's the one people assume rather than design. What happens after a reviewer objects, who picks it up, and how quickly, determines whether objecting is a sensible thing to do or a way of creating work for yourself.
Questions to Ask Before You Commit
On basis. What is the reviewer checking against? A bad answer is the output.
On visibility. Can they see what it worked from? A bad answer is that it's explainable.
On authority. Can they overturn a result? A bad answer is that they can escalate.
On granularity. Can they hold one item? A bad answer is that they can reject the batch.
On volume. How many per person per day? A bad answer is that it's manageable.
On placement. Does the reviewer know these people? A bad answer is that they're trained.
What Getting This Wrong Costs
The first cost is a control everybody believes in that isn't operating. Human review is the answer given whenever anybody raises a concern about these features, and once given it tends to close the conversation. If the reviewer can't actually judge what they're shown, the assurance was false, and every decision downstream rested on it, including the decision to use the feature more widely.
That last part is what compounds. A feature considered safe because it's reviewed gets extended to more cases, higher stakes and larger volumes, and each extension leans on the same assurance nobody has tested.
The second cost lands on the reviewer. Being asked to approve things you have no basis for judging is an uncomfortable position, and people in it generally know. They also carry the exposure: their approval is in the record, and if something goes wrong, the record points at them. That's an unfair position to put somebody in and it's usually accidental.
It tends to go unmentioned, too, because raising it sounds like complaining about your own competence. The people best placed to tell you that a review isn't working are the ones least comfortable saying so, which is why the question has to be asked of them directly.
The third cost is the approval record itself, which outlasts everybody's memory. A document saying a person checked this is produced, filed and relied on by people who weren't there and have no way to know what the check consisted of. When it turns out to have been a glance, the organisation has a documentation problem on top of whatever the original error was.
The cleanest way out of that is to record what was checked rather than that it was approved. A line naming what the reviewer verified is both more useful later and harder to produce without actually doing it.
So do three things. Write down, per feature, what the reviewer is checking, in one sentence somebody could act on. Make sure they can see what the output was based on, which is usually a setting. And find out whether any review has ever changed an outcome, because that single question tells you whether you have a control or a signature.
None of the three needs a vendor, a budget or a project, which is unusual for anything in this area. They're decisions about what people are shown and what they're allowed to do, and they sit entirely with whoever owns the process.
When You Are Ready to Go Further
Start by asking the person who does the reviewing what they actually do. Not the process owner, the reviewer. How long they spend, what they look at, what would make them refuse, and whether they ever have. The gap between that conversation and the documented process is usually where the whole problem lives, and it takes twenty minutes to find.
Ask it in a way that makes an honest answer safe. Somebody who suspects the real answer reflects badly on them will describe the documented process back to you, and that's the one conversation in this piece you can't afford to get a polite answer from.
Then fix what they're shown before fixing anything else. Seeing the source material alongside the output is frequently a configuration option, it converts guessing into checking, and it's the highest-return change available here. If volume makes source-checking impossible, that's a finding about volume rather than a reason to abandon it.
Check what they're shown about the question, too. An output answering something subtly different from what was asked is among the most common failures, and it's invisible unless the original request sits next to the result.
Finally, give them somewhere to go that isn't approve or halt. A way to query a single item, hold it, or route it to somebody who knows more, while everything else proceeds. Refusal has to be a realistic option or it isn't an option, and no amount of training substitutes for it.
Name who picks up a held item and how quickly. A hold that disappears into an unattended queue teaches reviewers within a fortnight that holding things is pointless, and after that the path exists on paper only.
HROpsLab publishes independent comparison work across HR tooling. We sell nothing, we take no vendor money, and we publish no paid placements. If the next step is understanding what your current tooling shows a reviewer, our comparison work is one place to start.
Frequently Asked Questions
What does human in the loop actually mean?
In practice it means a person sees an output before it takes effect, which is a much weaker guarantee than it sounds. What determines whether it's a control is what that person is shown, whether they hold knowledge the feature lacked, and whether they can realistically change the outcome. A reviewer handed an output with no source, no context and no way to refuse is performing an approval rather than a check, and the phrase covers both arrangements equally well, which is why it needs examining rather than accepting.
Does human review actually catch AI errors?
It catches the errors the reviewer is positioned to catch, which depends almost entirely on what they know and what they're shown. Somebody who knows the individual will catch a general statement misapplied to them. Somebody who knows what changed last month will catch a stale answer. Somebody shown only a fluent output with no source will catch very little, because confident wrongness reads exactly like confident correctness. The review is as good as the basis the reviewer has for disagreeing.
What does a reviewer need to see?
The material the output was based on, at minimum, because that converts the task from judging plausibility into checking a claim against a source. Also useful: the question as it was actually asked, since a correct answer to a misread question is a common failure nobody counts. Showing the source is frequently a configuration option rather than a development project, which makes it the highest-return change available in most setups, and it's worth asking the vendor whether it exists.
Why do approvals become a formality?
Because of arithmetic and options rather than character. A reviewer with more outputs than minutes develops a faster method within weeks, and that method is reading for plausibility. A reviewer whose only alternative to approving is halting a process everybody depends on will approve, because the cost of refusing falls on colleagues. Both are rational responses to how the review was designed, and neither is fixed by asking people to be more thorough, which is the usual response.
Is reviewing a sample of outputs enough?
It depends which job you want done, and these get confused. Sampling is a good way to learn how a feature behaves generally, to find patterns, and to decide whether to keep using it. It does nothing to prevent a wrong output reaching a specific person, because the one that reaches them is precisely the one you didn't sample. Both jobs are worth doing and they need different arrangements: sampling for measurement, a check before the output lands for prevention.
Who should review AI output in HR?
Whoever holds the knowledge the feature lacks, which for anything about an individual usually means somebody who knows that individual. This runs against the instinct to centralise review for consistency, and centralising removes the main thing a human contributes, since a reviewer who has never met the person has no way to know their arrangement is unusual. For outputs about policy or process, the right reviewer is whoever knows what your organisation actually does as opposed to what's documented.
What if the reviewer can't tell whether an output is right?
Then you don't have a review, and it's better to know that than to carry a record saying otherwise. The fix is usually one of three things: show them the source so they can check rather than judge, move the review to somebody with the relevant knowledge, or narrow what the feature is used for so its outputs fall within what somebody can verify. What doesn't work is asking the same person to try harder with the same information.
Does human review make AI safe to use?
Not by existing, and treating it as a box that has been ticked is how nominal review spreads. It contributes protection in proportion to what the reviewer knows, what they can see, and whether their disagreement changes anything. A review that never alters an outcome is either unnecessary or not functioning. Separately, whether oversight is required in your situation and what form it must take differ by jurisdiction and are changing, so establish your own position with local advice rather than assuming a review satisfies it.
The reviewer isn't there to be careful. They're there to know something it doesn't.