TL;DR
- The core decision: how to tell a real AI capability from a feature name, before you buy it.
- When doing nothing is right: when no task would change, whatever the claim turns out to be.
- What has to be true: you can check something observable, because you can't check the model.
- How the options split: by what evidence the claim rests on, from prepared demo to live trial.
- Decision rule: evaluate everything around the model, since the model itself is closed to you.
- Outcome to expect: fewer impressive answers, and a couple you can actually act on.
Every Answer Is Confident and None of It Is Checkable
You're forty minutes into a demo. The AI features look good. You've asked three questions about them and received three fluent answers, and you couldn't repeat any of those answers back to a colleague.
This isn't because you asked badly or because anybody lied. It's a structural problem with evaluating this particular kind of software. You cannot inspect the thing you're being sold. You can't read it, you can't test it properly in the room, and the person demonstrating it frequently doesn't know the details either, because they're presenting a capability their own organisation bought from somewhere else.
Worse, the demonstration itself is weak evidence. A demo is built from material chosen because it works. That isn't dishonesty, it's how anybody would build a demo, and it means a weak capability and a strong one look identical in the room. The difference between them only appears on data nobody curated, which is to say yours, later.
Best tools for AI HR Tools
So the reframe is this. Stop trying to evaluate the model. You can't, and pretending otherwise produces the forty minutes you just spent. Evaluate everything around it instead: what it was shown, what it does when it isn't sure, what happens when it's wrong, who finds out, and what you're allowed to check before committing. All of those are observable, none of them need technical knowledge, and the quality of the answers tells you a great deal.
One boundary before we start. Whether a particular use of automated processing about people is permitted where you operate, and what you owe anybody in the way of information or assessment, differs by jurisdiction, differs by sector, and is changing quickly. Nothing here tells you what applies to you. Establish it with local advice. The governance side of this lives in the AI in the workplace material rather than here.
When You Genuinely Do Not Need to Act Yet
No task would change. Whatever the claim turns out to be, nobody's week is different. Then it doesn't matter whether the capability is real, and the evaluation can stop.
The claim is about something somebody spends time on. A person does the work by hand and the feature says it does that work. Worth interrogating properly, which is what the rest of this is about.
You're being asked to commit before you can check. A decision is wanted at a point where no evidence beyond the demo exists. That's a negotiating position rather than a constraint, and it's worth naming as one.
The edge case that forces it. Somebody has already bought it, or a team is using something nobody evaluated, and you're working backwards from a decision already made. The same questions apply, they just arrive later.
Five Questions This Reader Asks at 11pm
Can you evaluate an AI feature at all? Not directly, and accepting that early saves a lot of frustration. What you can evaluate is what surrounds it: the data it was shown, its behaviour when uncertain, the correction path, the change policy, and what you're permitted to test. That's a genuine evaluation, it just isn't the one people expect to be doing.
Does the demo prove anything? Very little. It proves the capability can work on material selected to make it work, which is the weakest useful claim available. It isn't evidence of nothing, because an unimpressive demo is informative. But a successful one tells you almost nothing about your own situation.
What's the single best question to ask? What it does when it isn't sure. Almost everything follows from the answer. A feature that produces the same confident output whether or not it has any basis is a different purchase from one that can say it doesn't know, and vendors rarely volunteer which they're selling.
Should you ask about the underlying technology? You can, and the answer will be less useful than you hope. Knowing what something is built on doesn't tell you how well this implementation performs on your data, and it tends to move the conversation somewhere neither of you can evaluate. Stay on observable behaviour.
What if they won't let you test on your own data? That's information. It might be reasonable, and there are legitimate reasons around handling your records. But an outright refusal with no alternative, no sample process and no trial period tells you something about how confident they are, and it's worth weighing as evidence rather than as an inconvenience.
What You Can Actually Check
| What you cannot verify | The observable thing that stands in for it | What a weak answer sounds like |
|---|---|---|
| How good the model is | What it does on your own awkward cases | It performs very well |
| What it was trained on | What they'll tell you about sources, plainly | Proprietary data |
| Whether it's confident | Whether it can decline or flag uncertainty | It always returns a result |
| How often it's wrong | What a wrong one looks like, and who sees it | It's highly accurate |
| Whether it will stay the same | Their policy on changes and notice | It's continuously improving |
| Whether your data is safe | What leaves, what's retained, what it trains | Enterprise grade security |
| Whether it explains itself | Whether you can see what it worked from | It's explainable |
| Whether it suits you | A trial on live data with a person checking | Customers love it |
The third row is the one to lead with. A feature that always produces an answer, regardless of whether it has any basis for one, has shifted every uncertainty onto whoever reads the output, and they usually can't tell. A feature that can decline, flag or express a range is handing you information you can act on.
The fifth row is asked almost never and matters more each year. If the behaviour can move without your knowledge, then every check you do at purchase is a check on a version you may not keep. The answer you want isn't that nothing changes; it's that you'll be told, and ideally that you can compare against what you had.
The last row is where the actual decision lives. Everything above it is diagnostic. A trial on live data, with somebody checking every output against what they'd have done themselves, is the only evidence that reflects your situation, and it's worth negotiating hard for.
Five Diagnostic Questions You Can Self-Assess Against
What would you need to see to be convinced? Decide before the next conversation, because deciding afterwards means deciding against whatever you were shown. Writing it down in advance is the single most useful thing on this list.
Which of your cases are awkward? Name three or four real ones: the unusual document, the person with the odd arrangement, the request phrased badly. These are your test set, and they're worth assembling once and reusing.
Who on your side can tell whether an output is right? If nobody can, no trial will help, because you'll be looking at outputs with no way to judge them. That person needs identifying before any test, not during it.
What happens if you're wrong about this? How much does the decision cost to reverse? Switching a bundled feature off is trivial. A separate tool with a year's commitment and a migration is not, and the answer should change how much evidence you require.
Are you evaluating the capability or the product? They're different. A capability can be genuinely impressive inside a product that doesn't fit your process, and the demo will show you the first while you're deciding about the second.
Work through these five with whoever will actually use the feature rather than whoever is running the purchase. The user knows which cases are awkward and what a wrong output would cost them, and those two answers shape everything else on this list.
Six Ways an AI Claim Gets Made, Reviewed
A demo on the vendor's own prepared data
They show you the feature working on material they chose. It earns its place as a starting point, because it shows you the interface, the shape of the output and what the experience is meant to feel like, and an unconvincing demo is a genuine signal.
Where it falls short is as evidence about you. The material was selected because it works, the cases were rehearsed, and anything that breaks the feature simply isn't in the deck. A weak capability and a strong one look the same here.
Watch it, then treat it as an illustration rather than a result. The useful move afterwards is asking what it does on something they didn't choose.
There's one thing a prepared demo does tell you reliably, which is what the output looks like in the interface. How a result is presented, whether it shows what it was based on, whether a person can see enough to disagree with it: all of that is real and visible, and it matters as much as the quality of the underlying capability.
Watch for the cases that get skipped over quickly, too. A presenter moving briskly past a particular screen is usually moving past something, and asking to go back is both reasonable and informative.
A demo on a sample of your data
You supply material and they run the feature on it in front of you. It earns its place as a genuine step up: the data is yours, the cases are real, and the failures that appear are failures that would appear for you.
Where it falls short is preparation and selection. There's usually a gap between supplying the sample and seeing the result, and what happens in that gap isn't visible to you. The sample is also small, so anything rare in your data won't be represented.
Ask for it, and include your awkward cases deliberately rather than a clean representative set. A clean sample tells you what you already assumed.
Establish what happens to the sample as well. You're handing over real records about real people, and where they go, how long they're kept and whether anything is retained afterwards are reasonable questions to ask before supplying anything. Obligations here differ by jurisdiction and sector, so take local advice rather than assuming a standard arrangement.
Ask to see the failures too. A vendor who shows you which of your cases the feature handled badly, and can explain why, is demonstrating something more useful than a clean run.
A trial on your live data with a person checking every output
You run it properly for a period, and somebody who knows the work compares every output against what they'd have done. It earns its place as the only evidence that actually reflects your situation, across your real distribution of cases.
Where it falls short is cost. It's real work for a real person over a real period, it happens alongside their existing job, and the checking is the expensive part rather than the running. Organisations frequently start one and quietly stop checking halfway through, which converts it into a period of unevaluated use.
The thing to push for, and to resource honestly. Agree in advance who checks, how many, and what would count as a bad result.
The last of those is what makes the trial capable of producing a no. Without a stated threshold, the result is a pile of outputs and a general impression, and general impressions after weeks of invested effort tend to be favourable regardless of what happened.
Run it long enough to include something unusual. A fortnight of ordinary cases exercises the part of the distribution that was never in doubt, and the interesting behaviour lives at the edges where your real work occasionally goes.
A stated capability with no demonstration at all
It appears on a feature list or in an answer, and nothing is shown. It earns its place only as a claim to be tested later, which is to say it earns very little.
Where it falls short is obvious and worth stating anyway, because these claims carry surprising weight in write-ups and business cases. A capability nobody has seen work has exactly the evidential status of a capability nobody has seen work, however reputable the source.
Write it down as unverified. If it matters to the decision, make seeing it a condition rather than a follow-up.
The reason to be firm about this is that unverified claims don't stay labelled. They migrate into summaries, business cases and comparison tables, where they sit alongside things you actually watched work, formatted identically and carrying the same weight. Nobody does this deliberately; it happens because the distinction lives only in somebody's memory of the meeting.
A claim about what the feature will do in a future release
The capability isn't there yet, and the timeline is described with confidence. It earns its place where you have a long horizon and the vendor has a track record you can actually check with existing customers.
Where it falls short is that roadmaps are intentions, priorities move, and the version you're buying is the one that exists today. The claim also tends to arrive precisely when a current gap has been identified, which is worth noticing.
Buy what exists. Treat anything forthcoming as a bonus rather than as part of the evaluation, and say so plainly in your notes.
If a forthcoming capability genuinely is the reason for the purchase, that's worth saying out loud, because it changes the conversation. It also tends to change the commercial terms, since a vendor who knows you're buying on a promise has an interest in the promise being contractual rather than conversational.
A claim resting on a technology name rather than a result
The answer describes what the feature is built on and stops there, as though the name settles the question. It earns its place as an honest answer to a question about architecture, which is occasionally what you asked.
Where it falls short is that it answers a question you shouldn't be asking. Two products built on the same foundation can perform completely differently on your data, because almost everything that determines quality happens in how the capability was applied rather than in what it started from.
Redirect. Ask what it produces and what a wrong output looks like, and note whether the conversation can get there.
If it can't, that's the finding rather than a dead end. A vendor unable to describe what their feature outputs, in plain words, on request, is telling you something about how well they understand what they've assembled, and that matters more than which foundation it rests on.
Some will route the question to somebody else who can answer it, which is the best response available and worth treating as a positive signal rather than an inconvenience.
The Decision Table
| Situation | Scale | Setup | Primary Pain | Recommended Starting Point |
|---|---|---|---|---|
| No task would change | Any | Any | None | Stop evaluating |
| Only a prepared demo offered | Any | Evaluating | No evidence about you | Ask for your own awkward cases |
| Sample run available | Any | Evaluating | Small and curated | Supply the difficult ones deliberately |
| Trial possible but unresourced | Any | Evaluating | Checking is the real cost | Name the checker before starting |
| Nobody can judge an output | Any | Evaluating | No trial will help | Identify that person first |
| Capability claimed, never shown | Any | Evaluating | Unverified carries weight anyway | Record it as unverified |
| Promised in a future release | Any | Evaluating | Buying an intention | Evaluate what exists today |
| Answer names a technology | Any | Evaluating | Unanswerable ground | Redirect to output and failure |
| Reversal would be expensive | Any | Committing | Evidence bar is too low | Raise the bar to match the cost |
The fifth row stops the whole exercise and is worth checking before anything else. Trials produce outputs, somebody has to judge them, and if nobody in your organisation can say whether a given output is right, you'll finish the trial with a pile of results and no verdict.
The ninth row is the calibration everybody skips. The amount of evidence you should demand depends on what reversing the decision costs, and those two are routinely mismatched in both directions: exhaustive evaluation of a feature you could switch off tomorrow, and a demo-based decision on something with a year's commitment behind it.
The second row is the most common starting position and it isn't a problem in itself. A prepared demo is a reasonable first meeting. It becomes a problem only when nothing else ever happens, and the way that occurs is drift rather than decision: the trial gets deferred, the timeline compresses, and the demo quietly becomes the evidence base by default.
The Demo Is Built to Work
Selection is the whole problem. The material in a demo was chosen because it demonstrates well. Anything that breaks the feature was removed long before you saw it, not to deceive you but because nobody builds a demo out of their failures.
Success carries less information than failure. A demo that goes badly tells you something real. One that goes well is consistent with almost any level of underlying capability, which makes it close to uninformative about the question you're actually asking.
Nothing here assumes bad faith. Most vendors aren't overstating deliberately, and plenty are describing a capability they bought from somewhere else and don't fully understand either. The questions work because they're specific, not because they're suspicious.
Your awkward cases are the test. Assemble three or four genuinely difficult real examples and reuse them across every vendor conversation. It's the cheapest evaluation asset you can build, it makes products comparable, and it takes an afternoon once.
Watch what happens at the edge, not the centre. Everyone handles the ordinary case. Ask what the feature does with something incomplete, something contradictory, something it hasn't seen before. The answer, and whether anybody can answer, is the signal.
The interface is evidence too. Whether an output shows what it was based on, and whether a person looking at it has enough to disagree, is visible in the demo and matters as much as the quality behind it.
Ask what it does when it isn't sure. Then ask to see that. A feature that can decline or flag uncertainty is handing you something you can build a process on. One that always answers has moved the uncertainty onto a person who frequently can't detect it.
Who is presenting matters. The person demonstrating frequently doesn't build the capability and sometimes doesn't know what sits behind it. That's not a criticism of them; it means some questions need routing to somebody else, and a vendor willing to do that is telling you something good.
Ask the same question twice, weeks apart. Not to catch anybody out. A capability that's genuinely understood inside the vendor gets described consistently by different people at different times, and one that isn't produces two different answers, which is worth knowing before you commit.
The reason to hold all of this loosely is that none of it is adversarial. Most vendors aren't overstating deliberately, and many are describing something they bought and don't fully understand either. The questions work because they're specific, not because they're hostile.
Where These Evaluations Go Wrong
| The failure | How it shows up | What would have to change |
|---|---|---|
| Demo treated as evidence | A decision resting on a rehearsal | Your own awkward cases |
| Clean sample supplied | It works, and tells you nothing | Deliberately difficult examples |
| Trial run without checking | A period of unevaluated use | A named checker, resourced |
| Nobody can judge outputs | Results with no verdict | Identify the judge first |
| Unverified claim carries weight | A capability nobody saw, in the case | Mark it unverified in writing |
| Evidence bar unrelated to cost | Over-evaluating the trivial, under the serious | Match the bar to the undo cost |
The third row is the most wasteful outcome available here, because the effort was actually spent. A trial where the checking lapses halfway looks like diligence, produces a favourable impression, and rests on nothing, which is worse than not running one.
The sixth row is the cheapest to fix and among the most consequential. Deciding how much evidence a decision warrants is a two-minute conversation about what unwinding it would take, and it prevents both the wasted scrutiny and the expensive shortcut.
The fifth row is a documentation failure rather than a judgement failure. People know perfectly well which claims they saw demonstrated, at the time. Three weeks later, in a written comparison, the verified and the unverified sit in the same table looking equally solid.
What to Put in Writing
| Artefact | Who owns it | When it is written | What it prevents |
|---|---|---|---|
| What would convince you | Whoever decides | Before the first conversation | Deciding against what you were shown |
| Your three or four awkward cases | Whoever knows the work | Before any demo | Clean samples that prove nothing |
| Verified versus claimed, per feature | Whoever evaluates | As you go | Unverified claims gaining weight |
| Who checks trial outputs, and how many | Whoever decides | Before the trial starts | Unevaluated use |
| What the vendor said about change | Whoever evaluates | At the time | Checking a version you won't keep |
| What applies to you, per jurisdiction | You, with local advice | Before committing | A position you assumed |
The first row is worth more than the rest combined, and almost nobody does it. Criteria written before you've seen anything are criteria; criteria written afterwards are a description of what you saw. The difference determines whether the evaluation could ever have produced a no.
The fifth row is the one that ages badly. What a vendor said about changes to behaviour is easy to establish during a sales conversation and impossible to reconstruct a year later, when the results have moved and somebody is trying to work out whether you were told anything.
Questions to Ask Before You Commit
On uncertainty. What does it do when it isn't sure? A bad answer is that it always returns a result.
On failure. What does a wrong output look like? A bad answer is that it's highly accurate.
On data. What does it read, and what leaves us? A bad answer is enterprise grade.
On change. What can move without us knowing? A bad answer is that it keeps improving.
On testing. What can we run on our own data? A bad answer is that the demo covers it.
On explanation. Can we see what it worked from? A bad answer is that it's explainable.
What Getting This Wrong Costs
The first cost is a decision made on a rehearsal. A demo built from selected material, presented by somebody who doesn't know what sits behind it, becomes the basis for a purchase, and the gap between that and your actual data only appears months later when the feature is embedded in somebody's week. Nothing improper happened. The evidence was just far weaker than it felt at the time.
The second cost is the trial that proved nothing. These are expensive, they consume a knowledgeable person's time for weeks, and they routinely end without a verdict because the checking lapsed or because nobody had decided in advance what a bad result would look like. That's the worst outcome in this whole piece: the full price of evaluation, and no evidence at the end of it.
It also makes the next evaluation harder. A team that has run one inconclusive trial is markedly less willing to run another, so the failure costs you the method as well as the answer.
The third cost is evaluating at the wrong intensity. Organisations reliably over-examine features they could switch off in an afternoon and under-examine commitments that take a year to unwind, because the intensity follows how interesting the feature is rather than how expensive the mistake would be. Matching the bar to the undo cost is free and fixes most of it.
So before the next conversation, do three things. Write down what would convince you. Assemble three or four genuinely awkward real cases. And name the person who could tell whether an output is right, because without them nothing else here works.
All three are yours to do and none of them depend on a vendor cooperating, which is what makes them worth starting with. They also carry across every conversation you have in this space, so the afternoon spent on them is spent once.
When You Are Ready to Go Further
Start with the awkward cases, because they're reusable and they cost an afternoon. Three or four real examples that are genuinely hard: the unusual document, the person with the odd arrangement, the badly phrased request. Run every vendor against the same set and the conversations become comparable for the first time.
Then push for a trial on live data and resource the checking properly. Name who checks, agree how many outputs, and decide in advance what would count as a bad result. That last part is what makes the trial capable of producing a no, and a trial that can't produce a no isn't an evaluation.
Give the checker somewhere to record what they saw as they go. A running note of which outputs were wrong and in what way is worth far more at the end than a remembered impression, and it takes seconds per case while the case is in front of them.
Finally, keep a plain record of which claims you saw demonstrated and which you were simply told. Both belong in the decision, and they belong in it differently. Writing the distinction down at the time is trivial, and reconstructing it weeks later is impossible.
The same record answers a question that arrives later and is otherwise unanswerable: what you were told about how the feature would behave, at the point you agreed to buy it. When the behaviour moves a year on, that note is the only thing standing between a conversation and a disagreement.
HROpsLab publishes independent comparison work across HR tooling. We sell nothing, we take no vendor money, and we publish no paid placements. If the next step is seeing how the tools in this space describe themselves, our comparison work is one place to start.
Frequently Asked Questions
How do you evaluate an AI feature in HR software?
Not by evaluating the model, which is closed to you, but by evaluating everything around it. What it was shown, what it does when it isn't sure, what a wrong output looks like and who would notice, what can change without your knowledge, and what you're permitted to test before committing. All of those are observable and none require technical knowledge. The answers also vary enormously between products that sound identical on a feature list, which is what makes the exercise worth doing at all.
What should you ask a vendor about their AI?
Lead with what the feature does when it isn't sure, because most of the rest follows from it. A capability that always returns a result has moved every uncertainty onto whoever reads the output, and that person usually can't tell a confident wrong answer from a right one. Then ask what a wrong output looks like in practice, what the feature reads to produce an answer, and what can change without you being told. Answers that describe benefits rather than behaviour haven't answered.
Does a product demo prove an AI feature works?
Barely. A demo runs on material chosen because it demonstrates well, which isn't dishonesty so much as how anybody builds a demo, and it means a weak capability and a strong one look the same in the room. The asymmetry worth remembering is that a demo going badly is genuinely informative, while one going well is consistent with almost any level of underlying quality. Treat it as an illustration of the interface and the output shape, then ask what happens on something they didn't choose.
How do you test AI on your own data?
Assemble three or four genuinely awkward real cases first: an unusual document, a person with an odd arrangement, a request phrased misleadingly. A clean representative sample will pass everywhere and tell you nothing. Then push for a trial on live data where somebody who knows the work checks each output against what they'd have done themselves. The checking is the expensive part rather than the running, so name that person and agree the volume before starting, or the trial becomes a period of unevaluated use.
What does a good answer about training data sound like?
A plain one. Somebody able to say, in ordinary words, what the feature reads at the moment of a request, what it was built on before it encountered your organisation, whether anything you supply is used to improve it for other customers, and what's retained. You don't need technical depth and shouldn't push for it. What you're testing is whether anybody in the conversation can describe it clearly, because vagueness here frequently means the person presenting doesn't know, which is itself worth establishing.
Should you ask which model a tool is built on?
You can, and it'll tell you less than you expect. Two products built on the same foundation can behave completely differently on your data, because almost everything determining quality happens in how the capability was applied rather than in what it started from. The bigger problem is that the answer moves the conversation onto ground neither party can evaluate, which uses up the meeting. Redirect to what the feature produces and what a wrong one looks like.
What if a vendor won't let you test on your own data?
Treat it as information rather than an obstacle. There can be legitimate reasons, particularly around handling employee records, and a good vendor will offer something: a sample process, a limited trial, a reference customer with a comparable situation. What's telling is a flat refusal with no alternative offered, because it means your decision has to rest entirely on a rehearsed demonstration. Whether that's acceptable depends on how expensive the decision is to reverse.
How do you compare two AI claims fairly?
Run both against the same awkward cases, and record separately what you saw demonstrated and what you were merely told. Those two things end up side by side in a written comparison looking equally solid, and three weeks later nobody remembers which was which. Also write down what would convince you before either conversation happens, because criteria produced afterwards tend to describe whatever you were shown, and an evaluation that can't produce a no isn't comparing anything.
You can't check the model. You can check almost everything around it.