AI HR Tools 24 min read

AI That Predicts People

A prediction is a claim about a population applied to one person. Six things teams predict, what the score is and is not, and how a forecast quietly makes itself true.

Sarah Mitchell Sarah Mitchell 24 min read
AI That Predicts People

TL;DR

  • The core decision: what to do with a number that claims to say something about a person.
  • When doing nothing is right: when nobody would act differently knowing it.
  • What has to be true: the score changes where you look, never what you conclude.
  • How the options split: by whether the prediction is about a population or an individual.
  • Decision rule: a score is a question to ask, not a finding to act on.
  • Outcome to expect: better-directed attention, and no decisions resting on a number.

The Screen With Everybody's Name On It

Somebody shows you a list. Every employee, and beside each name a number: their flight risk. The people at the top are apparently likely to leave.

Your first reaction is that this is useful, and it might be. Your second, if you sit with it, is to wonder what the number actually means. Not how it was calculated, which nobody in the room can tell you anyway, but what it's claiming. Is it saying this person is going to leave? That people like this person tend to leave? That something in their record resembles something in the records of people who left?

Those are three different claims with three different weights, and the screen presents all of them as one number beside a name.

Here's the reframe that keeps this useful and safe. A prediction is a statement about a population applied to an individual, and the application is where the strength drains out. Something can be genuinely true about a pattern across many people and tell you very little about the specific person in front of you. The number is real. What it supports is narrower than how it reads.

Which leads to the only rule worth holding: the score tells you where to look, never what's true. Used to direct attention, it's valuable and low risk. Used as a finding about somebody, it produces decisions nobody could defend and, worse, changes how a person is treated on the basis of something that was never a fact about them.

One boundary. What may lawfully inform a decision about an individual, what you owe the person concerned, and what assessment may be required differ by jurisdiction and are changing. Nothing here tells you what applies to you. Establish it with local advice. The governance side, including monitoring and fairness, sits with the AI in the workplace material.

Free Weekly Briefing Stay ahead of what's changing in HR and people ops.

Join 4,200+ leaders getting practical insights every week — no fluff, just signal.

Join Free →

When You Genuinely Do Not Need to Act Yet

Nobody would act differently. The number exists, nobody's behaviour changes, and it's a screen somebody built. Fine to leave, and worth asking why it's there.

It's directing attention usefully. Somebody looks at a group they'd otherwise have overlooked and has a conversation. That's the good version and it needs nothing added.

It's informing decisions about individuals. A score is affecting how somebody is managed, developed or considered. At that point the gap between what the number supports and what it's being used for has become a real problem.

The edge case that forces it. Somebody was treated differently because of a score and found out. That's a conversation you can't have well, and it's the reason to settle the rules first.

Five Questions This Reader Asks at 11pm

What does a flight risk score actually tell you? That this person's record resembles the records of people who left, in whatever terms the system holds. It isn't a statement about their intentions, which it has no access to, and it doesn't know about the conversation they had with their partner last week or the offer they turned down.

Do these predictions work? At the population level they can be informative, in the sense that a group flagged as higher risk may show more departures than a group that isn't. At the individual level that's much weaker than it sounds, because a pattern that holds across many people is compatible with being wrong about most particular ones.

What should you do with a score? Look. Have a conversation, ask how somebody is finding things, notice a team where several scores moved. The score's job is to direct a limited amount of attention, and the conversation is where you actually learn something.

Should it affect decisions? This is where the care belongs. Using a prediction to decide about somebody's opportunities, development or future is making a decision about an individual on the basis of a population claim, and it raises questions whose answers differ by jurisdiction. Establish yours with local advice before it happens rather than after.

Should you tell people? Worth thinking about properly, and partly a legal question that differs by jurisdiction. The useful non-legal test is whether you'd be comfortable if somebody saw their own score and asked what it was used for. If the honest answer is no, that discomfort is telling you something about the use rather than about the disclosure.

What the Score Is and Is Not

How people read it What the number actually supports The harm in the gap
This person will leave Their record resembles those who did A conversation framed around a conclusion
This person is disengaged Something in the data pattern-matched A manager treating somebody as a problem
A high score means act now It means look, if you have capacity Urgency attached to a weak claim
A low score means no concern It means nothing matched, not that all is well The quiet person nobody checked on
The number is precise It's an ordering with a number on it Confidence the method doesn't support
It knows why It has correlation, not explanation Explanations invented to fit the score
It's objective It reflects what's in the records Whatever the records encode, repeated
It's about them It's about people whose records resemble theirs An individual treated as a category

The fourth row is the one nobody thinks about and it matters. A low score doesn't mean somebody is fine; it means nothing in their record matched a pattern. The person quietly unhappy who has told nobody and changed nothing will score low, and a team relying on the screen will look straight past them.

The sixth row is where a lot of practical harm starts. The system offers a score and no reason, so a manager supplies a reason, and the supplied reason tends to be unflattering and to stick. The person is now managed according to an explanation nobody checked.

The last row is the one to hold on to. Everything a prediction knows is drawn from other people. Applying it to an individual is a reasonable way to direct attention and an unreasonable way to conclude anything, and the whole practical difference sits there.

Five Diagnostic Questions You Can Self-Assess Against

What would somebody do differently because of this score? If the answer is have a conversation, good. If it's anything that affects the person's opportunities, that's the point to stop and take advice.

Can anybody say what the number means? Not how it was built. What it's claiming. If nobody can articulate it, nobody should be acting on it.

Who sees it? A score visible to a manager becomes part of how they see somebody. That's a bigger decision than it looks and it's usually made by default when access is configured.

Would you show the person? A clean test. If showing somebody their own score and explaining its use would be uncomfortable, examine the use.

What about the low scores? If nobody looks at them, the screen is directing attention towards a pattern and away from everybody who doesn't match it.

Run these five with a manager who actually uses the screen rather than with whoever commissioned it. Managers can tell you immediately what they do when a score is high, and that answer is the real policy whatever the documentation says.

Six Things Teams Predict, Reviewed

Who is likely to leave

The most common prediction in this area. It earns its place because attrition is expensive, notice is short, and knowing where to direct retention effort has real value.

Where it falls short is that it can only see what's in the records, and the things that actually cause people to leave are largely outside them: a manager, an offer, a personal circumstance, a conversation. It also can't see the person who has already decided and changed nothing about their behaviour.

Use it to choose who to talk to this month. Never to conclude that somebody is leaving, and never in a way that changes how they're treated.

The conversation is the point, and it's worth being clear with managers that the score isn't the agenda for it. Somebody who walks in already believing a person is leaving conducts a very different conversation from somebody checking how things are going, and the second one produces information.

It's also worth noticing which of your own actions the score is reading. Where it draws on things like activity or engagement signals, it's partly measuring how visible somebody is rather than how settled they are, and those come apart for people working quietly.

Who is likely to perform well in a role

Predicting future performance, typically for hiring or placement. It earns a place because the decision is genuinely hard and people making it unaided are working from thin evidence too.

Where it falls short is that it necessarily learns from your existing judgements about who performed well, which encodes whatever those judgements contained. It's also predicting something your own organisation measures inconsistently, which makes the target itself unstable.

This is the category where decisions about individuals are most likely, so it's the one to take advice on first. What's permitted differs by jurisdiction and is changing, and the fairness questions belong with the governance material.

The target itself deserves scrutiny before any of that. Predicting performance means predicting whatever your organisation recorded as performance, and if those records were produced by inconsistent processes, the thing being predicted is unstable in a way no amount of method fixes.

Ask what happens to somebody the system scores poorly. If the answer involves them being less likely to be considered, that's an individual decision resting on a population claim, and it's precisely the case to settle with advice beforehand.

Who is ready for a move

Predicting readiness for promotion or a different role. It earns its place by surfacing people who might not be visible, which is a genuine benefit in a large organisation where visibility is unevenly distributed.

Where it falls short is that readiness shows up in records only for people whose development has been documented, which is itself uneven. The people already known are the people with the fullest records, so it can reinforce existing visibility rather than correcting it.

Useful as a prompt to look at people you weren't thinking about. Check who it doesn't surface, because that's the more interesting list.

The visibility problem is worth stating plainly. People whose development has been documented, who have had managers that wrote things down, who worked in parts of the organisation that record more, all have fuller records. Readiness predicted from records is therefore partly predicting who was well documented.

The useful version runs alongside whatever you'd have done anyway. Where the system surfaces somebody nobody had considered, that's a genuine find; where it confirms the list you already had, it's told you nothing.

What headcount will be needed

Forecasting future requirements rather than individual outcomes. It earns its place as the safest thing on this list, because it's a population claim used as a population claim and nobody is affected individually.

Where it falls short is ordinary forecasting difficulty. It projects from what happened before, and it can't know about the decisions your organisation hasn't made yet, which are usually the ones that matter most.

Low risk and genuinely useful. Treat the output as one input to a planning conversation rather than as a number to plan against.

The limitation is ordinary and worth keeping in view: it projects forward from what happened, so it's least reliable exactly when your organisation is about to do something it hasn't done before. Those are also the moments when people most want a number.

Because nobody is affected individually, this is the category where you can be relaxed about the method and focused on the assumptions. Ask what it assumes stays constant, and whether any of those things are about to change.

Which populations are at risk of something

Group-level predictions: a team, a location, a job family. It earns its place because this is what these methods are actually good at, and a group-level signal is a reasonable prompt to go and find out what's happening.

Where it falls short is that a group finding invites individual conclusions. Once a team is flagged, its members start being seen through that lens, and the population claim quietly becomes a set of individual ones.

The best use of prediction in HR. Keep the finding at the group level and go and ask people, because the answer isn't in the data.

The discipline required is resisting the slide to individuals, which happens without anybody deciding on it. A team flagged as at risk becomes a team whose members are viewed through that flag, and within a few weeks the population finding has quietly become a set of individual assumptions.

Group size matters as well. A finding about a small team is close to a finding about the specific people in it, which removes most of the protection that working at population level was supposed to provide.

Who is likely to accept an offer

Predicting acceptance during hiring. It earns a place because it can direct effort in a process where effort is limited.

Where it falls short is the self-fulfilling risk, which is sharper here than anywhere else. A candidate predicted unlikely to accept may receive less attention, less follow-up and a less enthusiastic process, and then declines.

Use it to add effort where acceptance looks uncertain, never to withdraw effort. That inversion is the whole of using this one safely.

The reason it matters more here than elsewhere is speed. Hiring runs on a short cycle with visible outcomes, so a prediction that reduces effort gets confirmed quickly and looks accurate, which makes it very hard to argue against once it's embedded.

Worth checking what the prediction is drawing on, too. Where it reflects patterns in who accepted before, it can be reproducing something about your own process rather than about candidates.

The Decision Table

Situation Scale Setup Primary Pain Recommended Starting Point
Nobody acts on it Any Deployed None, but why is it there Ask what it's for
Directing conversations Any Deployed None Change nothing
Affecting individual decisions Any Deployed Population claim, individual use Stop, and take local advice
Managers see scores by default Any Deployed It becomes how they see somebody Decide access deliberately
Nobody can say what it means Any Deployed Acting on an unarticulated claim Write the claim in one sentence
Low scores never examined Any Deployed The quiet person nobody checked Look at the bottom too
Score with no reason attached Any Deployed Managers invent explanations Say plainly there's no reason given
Effort withdrawn from low scores Any Hiring Self-fulfilling Invert it, add effort instead
Person found out and asked Any Deployed A conversation you can't have well Settle disclosure, with advice

The third row is the line this whole piece is drawn around. Everything on the attention side is defensible and useful. Everything on the decision side needs a position you've established rather than assumed, because what's permitted differs by jurisdiction and is changing.

The drift across that line is gradual rather than decided. A score used for conversations becomes a score mentioned in a discussion about opportunities, and nobody notices the moment it started informing outcomes.

The seventh row is the practical fix that costs nothing. A score presented without any explanation invites managers to construct one, and constructed explanations about people tend to be unflattering and durable. An interface that says plainly that no reason is available prevents a surprising amount of harm.

The fourth row is the choice most likely to be made by nobody. Visibility gets configured during setup by whoever is doing the configuring, usually towards showing more, and it determines whether this becomes an attention tool or a lens through which individuals are permanently viewed.

The Prediction That Makes Itself True

Attention changes outcomes. Somebody flagged as likely to leave who then receives more attention, a better conversation and a development opportunity may well stay. That's the system working and it also means the prediction can't be evaluated on what happened next.

The inverse is the dangerous one. Where a low predicted likelihood leads to less effort, the prediction contributes to its own confirmation. This is most visible in offer acceptance, where reduced enthusiasm produces exactly the outcome that was predicted.

Managers act on labels. A score visible to somebody's manager becomes part of how that person is seen, in ways nobody intends and nobody can undo. Ambiguous behaviour gets read through the label, and the label was never a fact.

Labels outlast the data. A score changes as records change; the impression it created in somebody's mind does not, and nothing prompts a manager to revisit a view formed six months ago.

Invented explanations stick. Given a number with no reason, people generate reasons. Those reasons are guesses, they're frequently uncharitable, and they persist long after the score has changed.

The guess usually flatters the guesser. Explanations constructed for somebody else's score tend to locate the cause in that person rather than in anything the organisation did, which is the least useful place to look.

Nobody can tell whether it was right. A prediction that doesn't come to pass wasn't necessarily wrong: you may have intervened, or something changed. That means the usual way of judging a method isn't available, which is an argument for modesty about all of it.

Which makes vendor performance claims hard to read. Whatever was measured, it was measured in a setting where people were presumably acting on the outputs, and that acting is part of what produced the result.

The person is the only real source. Everything a prediction offers is an inference from records. What's actually happening with somebody is available by asking them, and a score that prompts that conversation has done its entire job.

And asking costs less than the system did. The uncomfortable arithmetic in this whole area is that a manager having regular conversations with their team produces better information than any score, and the score is frequently bought because those conversations aren't happening.

The practical position that follows is narrow and workable. Predictions are good at telling you where to spend a limited amount of attention, they're poor at supporting conclusions about individuals, and the distance between those two uses is where every problem in this area lives.

Holding that line is easier if it's written down before anybody sees a screen. Once a list of names with numbers exists, the pull towards treating it as a finding is strong, and arguing against it in the moment sounds like obstruction.

Where These Arrangements Go Wrong

The failure How it shows up What would have to change
Score read as a finding A decision resting on a population claim Use it to direct attention only
Manager sees it by default It becomes how somebody is perceived Decide access deliberately
No reason given, reason invented An unflattering story that sticks State plainly that no reason exists
Effort withdrawn from low predictions The prediction confirms itself Invert the response
Low scores never looked at The quiet person nobody checked on Examine the bottom of the list
Person discovers they were scored A conversation with no good version Settle disclosure in advance, with advice

The second row is a decision made by default in most deployments. Whether an individual's manager can see their score is configured during setup, usually towards more visibility, and it's one of the most consequential choices available here.

The sixth row is the one worth pre-empting. Somebody will eventually learn that a number about them exists, and the quality of that conversation depends entirely on whether you decided in advance what you'd say. Improvised, it goes badly.

The fifth row is the quiet one. Attention flows towards the top of any ranked list by design, and the person who scores low because nothing about them matched a pattern is precisely the person no process will ever surface.

What to Put in Writing

Artefact Who owns it When it is written What it prevents
What the score claims, in one sentence Whoever owns the system Before anybody sees it Acting on an unarticulated claim
What may and may not follow from it Whoever owns the process Before rollout Population claim, individual use
Who can see it, and why Whoever owns the system Before rollout Visibility decided by default
What you'd say if somebody asked Whoever owns the process Before rollout An improvised difficult conversation
How low scores are handled Whoever owns the process Before rollout Attention only ever going one way
What applies to you, per jurisdiction You, with local advice Before any individual use A position you assumed

The second row is the one that does the real work. A written line saying a score may prompt a conversation and may not inform a decision about opportunity is short, unambiguous, and the only thing standing between a useful attention tool and a decision nobody can defend.

The fourth row takes ten minutes and saves a bad afternoon. Deciding in advance what you'd tell somebody who asks about their own score means the answer exists before the question does, which is the only way that conversation goes reasonably.

Questions to Ask Before You Commit

On the claim. What does the number assert? A bad answer is a risk indicator.

On use. What may follow from a high score? A bad answer is that it's up to the manager.

On visibility. Who sees it? A bad answer is that it's role based.

On explanation. Does it say why? A bad answer is that it's explainable.

On the bottom. What happens to low scores? A bad answer is nothing.

On the person. What do we tell them? A bad answer is that they don't see it.

What Getting This Wrong Costs

The first cost is somebody being treated according to a number that was never about them. A manager who sees a score starts reading ambiguous behaviour through it, conversations shift, opportunities are weighed differently. None of it is deliberate and all of it is real, and the person on the other end experiences it as their manager's attitude changing for reasons they can't identify.

That last part is what makes it corrosive rather than merely unfair. Somebody who could see the reason could argue with it. Somebody experiencing an unexplained shift in how they're treated has nothing to push back against.

The second cost is a prediction that confirms itself. Reduced effort towards somebody predicted to leave, or a candidate predicted unlikely to accept, produces exactly the outcome forecast, and the screen then looks accurate. The system appears to be working precisely where it's doing the most damage, which makes it very hard to argue against internally.

That's also why apparent accuracy is weak evidence here. A method whose outputs change behaviour towards the people it describes cannot be judged by whether its forecasts came true, and the better it appears to be performing, the more carefully that claim deserves reading.

The third cost is a conversation with no good version. Somebody discovers a number about them exists, asks what it's for and what it's been used for, and every answer is unsatisfying unless you decided in advance. That moment is entirely predictable, it arrives eventually in any organisation doing this, and the only preparation is having settled the rules while it was still hypothetical.

It rarely stays with one person, either. Somebody who learns they were scored tells colleagues, and the version that circulates is shaped by however the first conversation went.

So do three things before anybody acts on a score. Write in one sentence what it actually claims. Write what may and may not follow from it. And decide who can see it, because that single configuration choice determines whether this becomes a way of directing attention or a way of seeing people.

All three are writing rather than building, and all three are much easier before a screen full of names exists than afterwards.

When You Are Ready to Go Further

Start by writing the claim down, because the exercise is clarifying and usually uncomfortable. One sentence saying what a high score asserts, in plain language, without technical hedging. If nobody in the organisation can produce that sentence, nobody should be acting on the number, and that's a finding rather than a failure.

Ask the vendor for their version of that sentence too, and compare. A supplier who can state plainly what their number claims is a different proposition from one who answers with a description of the method.

Then decide visibility deliberately rather than accepting the default. Who can see an individual's score is the choice that determines most of the downstream risk, because a manager who can see it will factor it in whatever anybody intends. Keeping scores at the group level, visible to people thinking about populations rather than individuals, keeps almost all the value and removes most of the harm.

That middle position is worth considering seriously rather than treating as a compromise. Most of the genuine value here is in noticing that something is happening in a team or a location, and none of that requires anybody to see a number against a name.

Finally, look at the bottom of the list occasionally. Everything about these tools directs attention towards people who match a pattern, and the person quietly unhappy who has changed nothing detectable is invisible to it. They're also frequently the one you'd most want to have spoken to, and no score is going to tell you that.

The broader point is that a ranked list is a device for allocating scarce attention, not a description of your workforce. Anybody relying on it as the second will systematically miss whoever the pattern doesn't fit.

HROpsLab publishes independent comparison work across HR tooling. We sell nothing, we take no vendor money, and we publish no paid placements. If the next step is understanding what your current tooling predicts, our comparison work is one place to start.


Frequently Asked Questions

What does predictive analytics in HR mean?

Producing a number or category expressing how likely something is about a person or a group: that somebody might leave, might suit a role, might be ready for a move, or that a population faces some risk. The important property is that every such number is derived from patterns across many people and then applied to one. That application is where the claim weakens considerably, which is why the same output can be genuinely informative about a group and close to meaningless about the individual it's displayed against.

What does a flight risk score actually tell you?

That this person's record resembles the records of people who previously left, in whatever terms the system holds. It has no access to intentions, conversations, offers or personal circumstances, which is where most decisions to leave actually originate. It also can't see somebody who has already decided and changed nothing detectable. Read as a prompt to have a conversation it's useful; read as a statement that somebody is leaving it's claiming something the method cannot support.

Does attrition prediction actually work?

At population level it can be informative, in that a group flagged as higher risk may show more departures than one that isn't. At the individual level that's a much weaker claim than it appears, because a pattern holding across many people is entirely compatible with being wrong about most particular ones. There's also a measurement problem: a prediction that doesn't come true may have been wrong, or you may have intervened, so the usual way of judging whether a method works isn't cleanly available.

What should you do with a prediction about an employee?

Have a conversation. That's genuinely the answer, and it's not a dodge. The score's useful function is directing a limited amount of attention towards someone you might not otherwise have thought about, and everything you actually learn comes from asking them how things are going. What shouldn't follow is any change in how that person is managed, developed or considered, because that means acting on a population claim as though it were a finding about them.

Should you tell someone they've been scored?

Partly a legal question whose answer differs by jurisdiction and is changing, so establish your own position with local advice. Beyond that, a useful test is whether you'd be comfortable if somebody saw their own score and asked what it had been used for. Discomfort with that conversation usually indicates something about the use rather than about the disclosure. It's also worth deciding what you'd say in advance, because somebody will eventually ask and an improvised answer tends to go badly.

Should predictions influence decisions about people?

This is the line where care belongs, and it's the point to take advice rather than judgement. Using a prediction to decide about somebody's opportunities, development or future means applying a claim about a population to an individual, and what's permitted in that situation differs by jurisdiction and is changing. Establish your position before anybody does it rather than afterwards. Directing attention is defensible and useful; deciding outcomes is a different activity with different requirements attached.

Why do predictions become self-fulfilling?

Because they change behaviour towards the person they describe. Somebody flagged as likely to leave who receives more attention and a better conversation may stay, which is the system working well and also means it can't be evaluated on the outcome. The damaging version is the inverse: where a low prediction leads to less effort, the prediction contributes to its own confirmation. That's clearest with offer acceptance, where reduced enthusiasm produces exactly the decline that was forecast, after which the screen looks accurate.

What should you ask a vendor about a prediction feature?

What the number claims, in one plain sentence. Whether it offers any reason alongside the score, since a bare number invites managers to invent explanations that stick. Who can see an individual's score, because that's configured at setup and determines most of the downstream risk. What happens at the bottom of the distribution, since attention flows only one way by default. And what the feature does when it has little to go on, because a confident score built on a thin record is worse than no score at all.

The score tells you where to look. It never tells you what's true.

Share on X Share on LinkedIn

What to do next?

Explore More Articles

Dig deeper into HR Ops strategy, tools, and workflows built for real teams.

Browse the blog →
Join the HROpsLab Community

Connect with People Ops practitioners sharing real workflows, tools, and challenges.

Join now →