TL;DR
- The core decision: whether anybody in your process could detect a wrong AI output.
- When doing nothing is right: when the output only ever reaches somebody who'd notice.
- What has to be true: you can describe what a wrong one looks like, per feature.
- How the options split: by whether the error announces itself or looks exactly like success.
- Decision rule: ask what a wrong output looks like, never how accurate it is.
- Outcome to expect: the same features, with somebody positioned to catch them.
It's Highly Accurate
You ask how often the feature gets it wrong. You're told it's highly accurate. Everybody nods, and the meeting moves on, and nobody has learned anything.
The problem isn't that the answer is evasive. It might be perfectly sincere. The problem is that accuracy is the wrong question, and it's the wrong question in a way that feels like the right one, which is why it gets asked every time.
Here's why it doesn't help. Suppose a feature is right almost always. That sounds excellent until you ask what the remaining cases look like, and the answer for most of these features is: exactly like the successful ones. Same format, same confidence, same absence of any flag. A wrong output doesn't arrive looking wrong. It arrives looking like an output.
Best tools for AI HR Tools
Compare that with a feature that's right less often but fails loudly, throwing an error or returning nothing or flagging that it wasn't sure. The second is frequently the better purchase, because its errors get caught and the first one's don't. An error you can see costs you a minute. An error that looks like a result costs you whatever was decided on the strength of it.
So the reframe is to stop asking about frequency and start asking about appearance and detection. What does a wrong one look like? Who sees the output? Would that person be able to tell? And what happens between the output appearing and somebody acting on it? Those four questions are answerable, they need no technical knowledge, and they tell you what accuracy never could.
One boundary. Systematic error affecting a group of people raises questions of fairness, assessment and disclosure whose answers differ by jurisdiction and are changing. Those belong with the governance material we publish on AI in the workplace, and with local advice. This piece stays on detection: whether anybody would notice, and how.
When You Genuinely Do Not Need to Act Yet
The output only reaches somebody who'd notice. A knowledgeable person sees every output and would spot a wrong one. That's the arrangement you want, and it's worth recognising when you have it.
Somebody sees the output but couldn't judge it. The check exists on paper. Whether it's a check depends entirely on whether the person has any basis for disagreeing, and frequently they don't.
Nobody looks at all. The output flows into a record, a decision or a message without a person in between. At that point wrongness is reaching people, and the only question is how long before somebody notices.
The edge case that forces it. An error surfaced, and tracing it back nobody can say how long it had been happening. That's the signal that detection was never designed, and the fix is bigger than the incident.
Five Questions This Reader Asks at 11pm
How accurate are these things actually? Unanswerable, and worth abandoning. No figure anybody gives you is grounded in your data, your cases or your definition of correct, and a number quoted about somebody else's situation tells you nothing about yours. Ask what a wrong one looks like instead. That question has a real answer.
Why does it sound so confident when it's wrong? Because fluency and correctness are unrelated properties, and most of these features produce output in a consistent format regardless of whether they had any basis for it. The presentation is the same either way, so confidence in the output is a property of the interface rather than a signal about the content.
Would we even know? For most features, honestly, no, unless somebody deliberately set out to find out. That's the finding worth acting on. Detection in this area doesn't happen by default; it happens because somebody designed for it.
Does it get worse over time? It can, in two directions. The feature can change on the vendor's side, and your own data can shift so that the same behaviour produces different results. Both are invisible without something stable to compare against.
Who's responsible when it's wrong? Practically, whoever the output reached, because from their perspective their employer told them something. Where responsibility sits formally, particularly where an output informs a decision about somebody, differs by jurisdiction and is changing. Establish yours with local advice.
What a Wrong Answer Looks Like
| The error | What the reader sees | Who's positioned to catch it |
|---|---|---|
| Confidently untrue | A normal output, fluent and specific | Only somebody who knows the subject |
| Right generally, wrong for one person | A reasonable statement, misapplied | Somebody who knows that person |
| Systematically wrong about a group | Each case looks individually fine | Nobody, without looking across cases |
| Out of date, and doesn't know | A confident answer from a stale source | Somebody who knows what changed |
| Fails on input it can't handle | Nothing, a blank, or a refusal | Anybody, this is the visible kind |
| Right about the data, wrong about the question | An accurate answer to something else | Whoever asked, if they reread it |
| Plausible detail that was never there | A specific that reads as authoritative | Somebody who checks the source |
| Omits something important | A complete-looking answer with a gap | Almost nobody |
The fifth row is the only good news in the table and it's worth valuing properly. A feature that fails visibly, returns nothing, or declines to answer has handed you the error rather than hiding it. Teams frequently treat this as a weakness. It's the opposite.
The third row is the one no individual review will ever find. Each case looks fine on its own, which is exactly what makes the pattern invisible, and seeing it requires deliberately looking across many outputs rather than at any one. The fairness and assessment questions that follow belong with governance and local advice.
The last row is the hardest of all, because an omission leaves nothing behind. A summary missing the important point reads as a summary. Nobody reviewing it has the original in their head, and the absent thing is absent from the review too.
Five Diagnostic Questions You Can Self-Assess Against
Per feature, what does a wrong output look like? Write it down in a sentence. If you can't, you don't know what you'd be looking for, which means you aren't looking.
Who sees each output, and could they tell? Two separate questions. Plenty of outputs are seen by somebody with no way to judge them, and that's a check in name only.
What sits between the output and somebody acting on it? Trace one. If the answer is nothing, then every error reaches its destination, and the feature's failure mode is your process's failure mode.
Have you ever found one? If no error has ever been detected, the likeliest explanation isn't that there haven't been any. It's that nothing would have found them.
The same reasoning applies to a feature everybody likes. Satisfaction with an output is a judgement about how it reads, not about whether it was right, and the two come apart precisely where it matters.
Could you tell if it got worse? Without something stable to compare against, behaviour can drift for a long time before anybody notices, and by then nobody can say when it started.
Ask these five of whoever actually reads the outputs day to day. They generally know which ones they'd struggle to judge, and that answer is more accurate than anybody's description of the process.
Six Ways an AI Feature Is Wrong, Reviewed
Confidently producing something that isn't true
It states something as fact that has no basis. It earns a place on this list because it's the failure people have heard of, and knowing it's possible is most of the protection.
Where it does damage is in specificity. The invented detail is usually the convincing part: a date, a figure, a clause, a name. Those read as evidence that the output is well grounded, when they're frequently the least grounded thing in it.
Catchable only by somebody who knows the subject. Where an output will be read by somebody who doesn't, that's the risk, and it's about the reader rather than the feature.
The practical defence is to check the specifics rather than the whole. Reading a paragraph for overall plausibility catches almost nothing, because plausibility is what it's best at. Picking out the two or three concrete claims and verifying those against a source catches most of it in under a minute.
It's also worth noticing where an output is unusually detailed. Precision that exceeds what the underlying records could support is a signal, and it's visible to anybody who knows how thin those records actually are.
Right in general and wrong about one specific person
The statement is reasonable about people broadly and wrong about the individual in front of you. It earns a place because it's the most common failure in HR specifically, where almost everything has exceptions attached to particular people.
Where it does damage is that nothing marks the exception. Somebody with an unusual contract, a variation nobody recorded, an arrangement agreed years ago: the feature has no way to know and no way to indicate that it might not.
Catchable by somebody who knows the person. That's an argument for keeping outputs about individuals in front of people who know them, rather than routing them centrally.
Centralising this kind of review feels efficient and removes the only thing that made it work. A person in a shared service reading an output about somebody they've never met has no way to know that the arrangement is unusual, and the output gives them no reason to suspect it.
The exceptions also cluster. Long-serving people, anybody who transferred in from an acquisition, anybody with a negotiated variation: these groups carry most of the unusual arrangements, and outputs about them deserve more scepticism than average.
Systematically wrong about a group
It's consistently off in the same direction for a particular set of people, and each individual case looks unremarkable. It earns a place as the failure that individual review structurally cannot find.
Where it does damage is invisibility plus accumulation. Because no single output looks wrong, no reviewer flags anything, and the pattern continues for as long as nobody looks across cases rather than at them.
Requires deliberately examining outputs in aggregate. The fairness, assessment and disclosure questions that follow from finding one differ by jurisdiction, belong with the governance material, and need local advice.
The aggregate look is a different exercise from the ongoing check and won't happen as a by-product of it. Somebody has to sit down with a body of outputs, group them, and ask whether the pattern differs between groups, which is a scheduled piece of work rather than a habit.
It also needs deciding in advance what you'd do if you found something, because that decision is much harder to make once a specific pattern is sitting in front of you.
Wrong because the situation changed
The answer was right for the world as the feature understood it, and something moved. It earns a place because it's inevitable rather than exceptional: policies change, structures change, people change roles.
Where it does damage is that nothing about a stale answer looks stale. It's delivered with the same confidence as a current one, and the feature has no way to indicate that its source hasn't been touched in a long time.
Catchable by somebody who knows what changed, which means it's worst immediately after any change. Worth being deliberately sceptical of outputs in the weeks after a reorganisation or a policy update.
The awkward part is that the people who know what changed are usually the ones who made the change, and they're rarely the ones reading the outputs. Telling whoever does read them that something moved, and when, costs nothing and closes most of the gap.
Document changes are the quiet version of this. A policy superseded but still sitting in the same place will be read by a feature exactly as though it were current, and nothing about the answer will indicate which version it came from.
Refusing or failing on input it can't handle
It returns nothing, errors, or declines. It earns its place as the failure mode you should want, because it's visible, immediate and unambiguous.
Where it falls short is only in feeling like a defect. A feature that refuses on a fifth of its inputs seems worse than one that answers everything, and is frequently better, because you know which fifth you're handling yourself rather than discovering later which fifth was wrong.
Treat a refusal as information. A vendor whose feature can decline is offering you something more useful than one whose feature always answers.
What matters is what your process does with a refusal. If a blank output quietly becomes a blank field, or if people learn to rephrase until something comes back, the visibility has been thrown away and you're back to the silent failure with extra steps.
Route refusals somewhere a person handles them. That way the feature's honesty about its own limits actually buys you something instead of creating a gap.
Right about the data and wrong about the question
The output correctly reflects what was asked and the asker meant something else. It earns its place because it's the most common of all and it's rarely counted as an error at all.
Where it does damage is in the absence of blame. The feature did what was asked, the person asked what they meant, and the mismatch is in the interpretation between them. Nobody is wrong, and the output is used anyway.
Catchable by the asker rereading what came back against what they wanted. The best protection is an output that shows what it understood the question to be.
This failure is badly under-counted because nobody involved did anything wrong, so it rarely gets logged as an error at all. The output was correct, the question was reasonable, and the mismatch sat between them, which means it never appears in any tally of how often the feature is wrong.
It's also more likely the less precise the question, which makes it worst for exactly the open-ended questions people most want to ask.
The Decision Table
| Situation | Scale | Setup | Primary Pain | Recommended Starting Point |
|---|---|---|---|---|
| Knowledgeable person sees every output | Any | Any | None | Change nothing |
| Output seen by somebody who can't judge | Any | Enabled | A check in name only | Give them the source, or move the check |
| Nothing between output and action | Any | Enabled | Every error arrives | Insert a person, or narrow the use |
| No error has ever been found | Any | Enabled | Nothing would have found one | Look deliberately, once |
| Output goes to somebody without context | Any | Enabled | Invented specifics read as fact | Route it to somebody who'd know |
| Outputs about individuals routed centrally | Any | Enabled | Exceptions invisible | Put them in front of people who know |
| Just reorganised or changed policy | Any | Enabled | Stale answers look current | Heightened scepticism for a period |
| Feature always returns something | Any | Enabled | Uncertainty pushed onto readers | Ask whether it can decline |
| Nothing stable to compare against | Any | Enabled | Drift invisible | Keep a fixed set of examples |
The fourth row is the one to act on immediately and it's counterintuitive. An absence of detected errors is usually evidence about your detection rather than about the feature. The check is to go looking once, deliberately, at a set of outputs, and see what turns up.
The second row is the most common arrangement in practice and the one people are most satisfied with, because it looks like a control. Somebody sees every output. Whether they could disagree with any of them is a different question and rarely asked.
The ninth row is the cheapest item in the table and the one that answers the question you'll be asked during an incident. A retained set of examples with their original outputs turns how long has this been happening from an unanswerable question into a morning's work.
The Error That Produces No Signal
Silence is the expensive failure. Every error in this area falls into one of two categories: the kind that announces itself and the kind that looks exactly like success. Only the second costs anything real, and it's the one that vendor conversations never cover.
Confidence is formatting. The output looks the same whether or not there was a basis for it. Reading fluency as reliability is the single most common mistake made with these features, and it's made by careful people because the presentation genuinely doesn't distinguish.
The specific detail is the risk. Invented material is most convincing when it's precise: a date, a figure, a reference. Those are the parts a reader treats as evidence of grounding, and the parts most worth checking against a source.
Omission leaves nothing to find. A missing element in a summary is invisible to anybody who doesn't have the original, and the reviewer usually doesn't. This is the failure with no detection strategy short of going back to the source.
Length is not completeness. A long, well-structured output feels thorough, and thoroughness of form says nothing about whether anything important was left out of the substance.
Aggregate errors need aggregate looking. Anything systematically wrong about a group will pass every individual review, because each case is individually unremarkable. Finding it means examining many outputs together, deliberately, as a separate exercise.
Nobody volunteers that they can't judge an output. A reviewer with no basis for disagreeing will usually approve rather than say so, because saying so sounds like admitting they don't understand their own job. Asking them directly, and making it safe to answer honestly, is the only way that gap surfaces.
Detection is designed or absent. No feature detects its own errors for you, no process catches them by default, and nobody stumbles across them reliably. If you didn't build the check, there isn't one, and the first indication will be somebody affected by the output.
The check has to be cheap enough to survive. A review that takes too long gets abbreviated within weeks and abandoned within months, whatever anybody intended. Two or three specifics verified against a source, every time, beats a thorough review that stops happening.
The practical consequence of all six is that the important question about any feature isn't how good it is. It's whether the arrangement around it would surface a bad output before that output did anything. That question is answerable in a sentence per feature, and most organisations have never asked it.
It also reorders what you should care about when choosing between products. A feature that's modestly capable and honest about its limits fits into a safe arrangement easily; a stronger one that always answers confidently needs far more built around it to reach the same place.
Where These Arrangements Go Wrong
| The failure | How it shows up | What would have to change |
|---|---|---|
| Accuracy asked instead of appearance | A number that means nothing for you | Ask what a wrong one looks like |
| Reviewer can't judge the output | Approval with no basis | Give them the source, or move the check |
| Fluency read as reliability | Confident wrongness accepted | Check the specifics against a source |
| Individual review only | Group-level patterns never found | Look across outputs, deliberately |
| No baseline kept | Drift invisible for a long time | A fixed set of examples, rerun |
| No errors ever found | Mistaken for a good sign | Go looking once, properly |
The first row is where most of this starts. The accuracy question feels rigorous, produces a confident answer, and closes the subject, which means the conversation that would have been useful never happens.
The fifth row is the one that matters more each year. Behaviour in these features can move without any change on your side, and without something fixed to compare against, there's no moment at which anybody can say results are different now. A small set of retained examples, rerun occasionally, is the cheapest instrument available.
The third row is the one that defeats careful people. Nobody reads a fluent, well-organised output sceptically by default, because everything about the way it's presented signals that it was produced by something that knew what it was doing.
What to Put in Writing
| Artefact | Who owns it | When it is written | What it prevents |
|---|---|---|---|
| What a wrong output looks like, per feature | Whoever uses the feature | Before relying on it | Looking for nothing in particular |
| Who sees each output, and can they judge | Whoever owns the system | Now | A check in name only |
| What sits between output and action | Whoever owns the process | Now | Every error arriving intact |
| A fixed set of examples, with results | Whoever owns the system | Before rollout | Drift nobody can date |
| The aggregate review, and when it happens | Whoever owns the system | Before rollout | Group patterns never found |
| What applies to you, per jurisdiction | You, with local advice | Before decisions about people | A position you assumed |
The fourth row is small, dull and the most useful thing here. A handful of real examples with their outputs recorded, rerun every so often, converts the unanswerable question of whether the feature has changed into a comparison anybody can do in an hour.
The first row is the one to write before anything else, because every other item on this list depends on it. You can't design a check, brief a reviewer or run an aggregate review without a sentence saying what you're looking for.
Questions to Ask Before You Commit
On appearance. What does a wrong output look like? A bad answer is that it's highly accurate.
On uncertainty. Can it decline or flag that it wasn't sure? A bad answer is that it always returns a result.
On source. Can we see what it worked from? A bad answer is that it's explainable.
On detection. How would we know it was wrong? A bad answer is that users report issues.
On drift. What would tell us the behaviour moved? A bad answer is that it keeps improving.
On aggregates. Can we review many outputs together? A bad answer is that each is reviewed.
What Getting This Wrong Costs
The first cost is a wrong output reaching a person. Somebody is told something about their pay, their entitlement, their record or their application, and it isn't true, and it came from their employer. The technical failure behind it might be trivial. What lands on them isn't a technical failure, it's being told something wrong by the organisation that employs them, and that's a trust cost rather than an accuracy one.
The correction rarely restores the position either. Being told something wrong and then told the right thing leaves a person less certain about the next answer they get, which is a cost that persists well after the individual case is closed.
The second cost is the error nobody can date. When something is eventually discovered, the first question is always how long it's been happening, and without a baseline or an aggregate review there's no answer. That turns a contained problem into an open-ended one, because every output the feature has produced is now suspect and nobody can narrow the range.
The scope of the remedy follows from that range. Not knowing when something started means either checking everything or checking nothing, and both are bad positions to be in while people are waiting for an answer.
The third cost is the review that was never a review. Somebody looked at every output, approved each one, and had no way to disagree with any of them, which means the organisation carried the cost of a control while getting none of its benefit. This is the most common arrangement in practice and the hardest to see from inside, because it looks exactly like a process working.
It's also demoralising for whoever is doing it, once they realise. Being asked to approve things you have no basis for judging is an uncomfortable position, and people in it generally know, even when nobody has said so.
So do three things, none of which need a vendor. Write a sentence per feature describing what a wrong output would look like. Check whether whoever sees each output could actually tell. And go looking once, deliberately, at a batch of outputs, because an absence of found errors is not an absence of errors.
When You Are Ready to Go Further
Start by going looking, once, properly. Take a batch of real outputs from a feature you rely on, have somebody knowledgeable work through them against what they'd have produced themselves, and see what turns up. It's a few hours, it needs nobody's permission, and it settles the question that nothing else can.
Pick outputs from a period you've already moved past, so nobody is deciding anything while they look. That separates the exercise from the day job and makes it far more likely to be done carefully.
Then build the baseline, because it costs almost nothing and answers the question you'll need later. A handful of real examples with their outputs recorded today, rerun every few months, tells you whether behaviour has moved and roughly when. Without it, the answer to how long this has been happening is always that nobody knows.
Keep the examples somewhere outside the system they came from. A baseline held inside the thing it's meant to check moves when that thing moves, which defeats the point of having one.
Finally, add one aggregate review to the calendar. Looking across many outputs rather than at individual ones is the only way to find anything systematic, it's a different exercise from the ongoing check, and it won't happen unless somebody owns it. What you do with a pattern you find is a governance question, and belongs with local advice.
Decide who owns it by name rather than by team. Reviews assigned to a function get scheduled, deferred and quietly dropped; reviews assigned to a person with a date attached usually happen.
HROpsLab publishes independent comparison work across HR tooling. We sell nothing, we take no vendor money, and we publish no paid placements. If the next step is understanding what your current tooling gives you here, our comparison work is one place to start.
Frequently Asked Questions
How accurate are AI features in HR software?
It's the wrong question, which is awkward because it feels like the right one. No figure anybody quotes is grounded in your data, your cases or your definition of a correct answer, so a number derived from somebody else's situation tells you nothing usable about yours. The question that does work is what a wrong output looks like when it appears, because for most of these features the answer is that it looks exactly like a right one, and that single fact matters more than any frequency.
Why does AI sound confident when it's wrong?
Because fluency and correctness are unrelated, and the output is formatted the same way regardless of whether there was any basis for it. Confidence in what you're reading is a property of the presentation rather than a signal about the content. The practical consequence is that the most convincing parts of a wrong output, the specific date or figure or reference, are frequently the least grounded parts, since those are exactly the details a reader treats as evidence that the whole thing is well founded.
What does an AI error look like in HR?
Several distinct things, and it's worth separating them. Something confidently untrue. Something true in general and wrong about one particular person, which is the most common in HR because nearly everything has exceptions attached to individuals. Something consistently off for a group, where each case looks individually fine. Something correct for a world that has since changed. Something that failed visibly. And something that answered a slightly different question from the one asked. Each is caught, if at all, by a different person.
How do you spot an AI mistake?
By deciding in advance what you're looking for, because looking generally finds nothing. Write a sentence per feature describing what a wrong output would look like, then check whether whoever sees that output has any basis for disagreeing with it. Add one review that looks across many outputs together, since anything systematic passes individual review by definition. And go looking deliberately once, because an absence of detected errors usually says more about your detection than about the feature.
Do accuracy figures from vendors mean anything?
Very little for your purposes. Whatever was measured, it was measured on data that isn't yours, against a definition of correct that may not match your own, in conditions that don't resemble your ordinary week. The figure isn't necessarily dishonest, it's just not transferable. What you can usefully ask instead is whether the feature can indicate that it wasn't sure, whether you can see what it worked from, and what happens on an input it can't handle, because those are observable and they're about your situation.
What happens when AI is wrong about one person?
This is the characteristic HR failure, and it's the one worth designing around. An output can be entirely reasonable as a general statement and wrong for the individual it's applied to, because that person has an unusual contract, a variation nobody recorded, or an arrangement agreed long ago. Nothing in the output marks the exception, since the feature has no way to know it exists. The practical defence is keeping outputs about individuals in front of people who know those individuals, rather than routing them somewhere central.
Do AI features get worse over time?
They can move, which isn't quite the same thing, and it happens in two directions. The capability can change on the vendor's side, sometimes without a visible release. And your own data can shift, so identical behaviour produces different results. Both are invisible unless you have something stable to compare against, which is why a small set of retained examples with their original outputs is worth keeping. Rerun them occasionally and you can answer a question that's otherwise unanswerable.
Who is responsible when an AI feature gets something wrong?
Practically, you are, from the point of view of whoever received the output, because what reached them was a statement from their employer regardless of what produced it. Where formal responsibility sits, particularly where an output informs a decision about somebody, differs by jurisdiction and is changing, so establish your own position with local advice rather than assuming a vendor arrangement transfers it. What's within your control either way is whether anybody would have caught it before it arrived.
Accuracy is unanswerable. What a wrong one looks like is not.