Measuring AI ROI in HR Without Trusting the Vendor Numbers

Time-to-fill drops when the market softens, not because your tool works. Six ways to measure AI in HR, what each one actually proves, and how to report an honest small gain.

James Carter James Carter 21 min read
Measuring AI ROI in HR Without Trusting the Vendor Numbers

TL;DR

  • The decision: Renew, renegotiate or walk away from an HR AI tool when only the vendor has measured its value.
  • When to do nothing: If the tool is fixing a problem you don't have, killing it's the cheapest ROI of all.
  • What has to be true: A baseline taken before rollout, plus a held-back comparison, plus a quality metric that reads after the time saved does.
  • How the options split: Pilot-first, control-group, or trust the dashboard, and only one of these survives a tough renewal.
  • Decision rule: If you can't name three things you measured yourself, you're not measuring ROI, you're reporting it.
  • Expected outcome: A small, defensible number on a slide, and a renewal conversation that doesn't end in a fist fight.

Priya is the head of HR operations at a 4,800-person logistics company. Tuesday morning, ten days before the budget review. She opens the vendor's quarterly business review pack. Time-to-fill is down. Ticket deflection is up. Recruiter hours saved is on a chart that goes up and to the right. Every number has a green arrow. The vendor wants a three-year renewal and a seat expansion.

She closes the pack and opens a blank slide. She has nothing to put on it that came from her own team. The market softened in the same quarter. Two of her recruiters went on parental leave. The ticket deflection could mean the chatbot is helpful, or it could mean people gave up and emailed a person instead, and the chatbot doesn't see that. None of this is in the deck.

But here's what Priya has come to realise in the last week, and it's the only thing that has changed how she is going into the room. The real issue isn't whether the tool works. The real issue is that she has no way to prove it either way, on a slide that the CFO won't pick apart. And her vendor knows that.

When You Do Not Need to Act on This Yet

There are four honest stages here, and the first one matters.

Your current setup is genuinely fine. A small HR team running a clean ATS and a decent intranet doesn't need an AI layer at all. Picture the HR manager at a specialist engineering firm where hiring runs at a trickle and every employee question reaches a named person the same day. Nothing there is bleeding. The cost of doing nothing is zero, and the cost of a bad rollout is a year of internal goodwill. If the trigger for evaluating AI was a vendor cold email or a conference talk, close the tab. Can you name the queue that's backing up? If not, you're building the case backwards from a demo, which is how HR functions end up owning software they quietly stop opening.

Friction, not failure. The tool is in place and it works, but recruiters still copy data between two systems, or the chatbot hands the hardest tickets to a human anyway. Think of a payroll coordinator who gets a tidy AI summary of every query, then opens the full case anyway, because the summary drops the one detail that decides it. Nothing here is broken enough to escalate, which is why this stage is the useful one. You can still capture what the work costs today, before somebody proposes a redesign and the old process disappears. Measurement is cheap while nobody's defending anything. Wait until renewal and you'll be rebuilding a baseline from the wreckage of a process the tool reshaped.

Free Weekly Briefing Stay ahead of what's changing in HR and people ops.

Join 4,200+ leaders getting practical insights every week — no fluff, just signal.

Join Free →

Real risk on the table. You're about to renew, expand, or roll out to a second region. Spend is committed. Headcount depends on the number. A wrong number is now a career problem, not an analytics problem. The tell is a sentence in your own draft business case that begins "the vendor reports". If somebody outside HR now signs off the cost, you're in this stage whether you feel ready or not. This is the audience this article is written for. The uncomfortable part is that most of the work you needed sits behind you, in a window that closed the day the tool went live.

The edge case. The tool is being procured for compliance reasons, not efficiency. Measuring ROI is the wrong frame. A works council wants documented consistency in shortlisting. An internal auditor wants to know how a scoring step gets recorded and who can override it. Nobody in that conversation is asking about recruiter hours. Build an ROI case here and you'll win an argument nobody was having. The conversation belongs to legal and audit, and you should be in that room, not running a pilot. Where the tool touches hiring decisions or employee records, the exposure is legal rather than financial, and its shape depends on where your people sit. Take local advice, not a view from a blog.

The 11pm Questions, Answered

"Is time saved the same as money saved?" Not always. If a recruiter banks ten hours a month and spends nine of them in more sourcing calls, you paid for the tool and didn't free a head. Saved time only turns into money where it collects somewhere a person can act on: a role you decide not to backfill, or a contract you stop renewing. Spread across eleven people in ninety-minute slivers, it's absorbed within a fortnight. It matters because this is the single most common error in vendor decks. Promise the saving anyway and finance will go looking for it at year end. They won't find it, and they will remember.

"Can I build a baseline after the fact?" Partially. You can reconstruct it from old tickets, archived requisitions, prior quarterly numbers. It will be rougher than a pre-rollout capture, but rougher is better than nothing. The trick is to be loud about the roughness rather than hide it. Name the window you pulled from, and let the room see the edges of your own data. A reconstructed baseline that states its weaknesses survives questioning. A polished one that can't explain its figures won't survive the first analyst who asks. It matters because most readers reading this missed the pre-rollout window.

"Why does the vendor benchmark look so different from mine?" Because the vendor is averaging across customers with different stacks, different process maturity, and different definitions of a successful ticket. The average also excludes the customers who left, so the cohort you're compared against is the cohort that stayed. Ask how many customers sit inside the figure. Ask what counts as resolved. It matters because a benchmark is a comparison to other people, not to where you were. Get it wrong and you'll spend the renewal defending a gap between two numbers that never measured the same thing.

"How long should a pilot run?" Long enough to see a full cycle of the work it touches, so usually a quarter for recruiting, a quarter for service. Less than that and you're measuring ramp-up, not steady state. The cycle that counts is the work's, not the calendar's. A requisition that takes weeks to close needs a window wide enough to hold several, opening to start date. A service desk only shows its shape after a peak, so a pilot that dodges the busy period tells you about a quiet month. It matters because pilots that end early produce great vendor case studies and bad decisions.

"What if my honest number is small?" A small, defensible gain is a good answer. A large, soft number isn't. Small numbers survive because they can be walked through line by line, and because whoever brought one has shown they'll report against their own interest. That buys standing the deck can't. It matters because the meeting you're trying to survive is the one where a soft number gets opened up. Inflate it and you don't merely lose this renewal. You lose the right to be believed on the next tool, and the next headcount request.

Three Categories the Approaches Split Into

Trust the vendor dashboard. Use the numbers the platform reports, the deflection rates, the time saved, the hiring velocity. You take the vendor's definitions along with their arithmetic, and in exchange you spend nothing. Right when the tool is young, the team is small, and the cost of independent measurement is more than the renewal at stake. A single-site charity paying a modest monthly fee for a policy assistant has no business running a control group. The measurement would cost more than the software. Fails when the spend gets large enough to need a board-level answer, because a CFO won't sign a renewal on someone else's spreadsheet. That failure has a visible shape. Somebody outside HR asks how a deflected ticket is defined. You say you'll come back to them, and the meeting moves on without your renewal in it. If the dashboard really is your only source, export it monthly and keep your own copy. Definitions change with product releases.

Pilot-first measurement. Stand up the tool in one team or one region, hold the rest back as a comparison, and run for a full cycle. The held-back group isn't a courtesy to sceptics. It's the only thing that tells you whether the market moved or the software did. Right when the rollout is optional and the renewal is more than a year away. A shared services lead handed a budget and no deadline is in the best position she'll ever be in, and probably doesn't know it. Fails when leadership has already committed publicly to the rollout, because the held-back group becomes a political problem. Once a chief executive has said the word on an all-hands call, the region that goes second is being punished, and its manager will say so. Watch for the softer version too. The pilot team gets a trainer and a weekly check-in wave two will never see, and what you've measured is the attention.

After-the-fact reconstruction. Pull historical data, build a synthetic baseline, and run a lighter comparison against a peer team. This is the salvage job, and it's what most readers are doing. Right when the tool is already in and the renewal is on the calendar. A benefits manager three weeks out from her renewal meeting has no other option, and careful salvage work is genuinely defensible when she shows the workings. Fails when the historical data is dirty or the tool has already changed the underlying process, so the "before" no longer exists in a clean form. If the intake form was rewritten the week the assistant went live, your old ticket volumes are counting a different thing, and no amount of care repairs that. One rule rescues this method more than any other. Decide what you're comparing before you look at the numbers. Decide afterwards and you'll choose the flattering window without noticing you chose it.

Five Diagnostic Questions

Do I have a baseline from before the tool was live? If the answer is no, every "after" number in your deck is unanchored. Don't answer this one from memory. Go looking for a dated export or a saved report from before go-live, then check whether its definitions still match today's. Most people find something. Far fewer find something comparable. This is the single most common failure and the easiest one to fix next time.

Can I name a team or region that didn't get the tool yet? A control group in HR rarely means randomisation. It usually means the second wave has not been switched on. Answer it with the rollout plan open, and look for a group that matches on the things that move your metric: similar roles, similar volume. If you can name it, you can use it. If you can't, you're running a before-and-after with a confounded baseline, and it's better to say so on your own slide than to have it said to you.

Do I know which downstream quality metric I am willing to wait for? First-year attrition of hires, quality of hire ratings, error rates in payroll or case handling. Pick it now, not in the renewal meeting. The way to choose is to ask what breakage would look like. If the tool were quietly making worse decisions faster, where would that surface, and who would notice before you did? That's your metric. You will need patience, because the read time is longer than your renewal cycle.

Is the licence fee the largest line on the total cost? It isn't. Implementation, integration, internal time, retraining, and the time your team spends running the tool all add up to several times the licence. To answer it honestly, count your own people's hours: the integration work, and the person who now owns the prompt library and the escalation rules. Put a rate on them. If your business case treats licence as the cost, your ROI is fiction, and finance will build the real figure for you at the least helpful moment.

Have I asked the vendor for the raw data behind their benchmark? The request is reasonable. Ask in writing, ask for the cohort size alongside the metric definition, and say when you need it. The answer will tell you almost everything you need to know about whether the renewal conversation is going to be a partnership or a fight. A vendor who sends you the methodology has one. A vendor who sends a case study instead has already answered you.

Six Ways to Measure, and What Each One Proves

Before-and-after on a single metric

Compare the same number before and after rollout. Time-to-fill in the quarters either side. Cost per case in the months before the assistant went live, then the months after. Cheap, fast, defensible. It earns its place because most HR teams can run it without asking permission, and because one clean metric is something a non-HR buyer can hold in their head while you explain the rest. Weak because the market, the team, and the season all change in the same window, so you can't tell what moved the number. A hiring freeze at your two biggest competitors will improve your time-to-fill more than any assistant will. So will a good run of referrals. The specific failure to watch: the metric moves in your favour, you present it, and somebody in the room already knows about the freeze. Reads in a quarter.

A held-back control group

Hold one team or region back for one cycle and compare. In an HR setting this almost never means randomisation. It usually means wave two, switched on a quarter after wave one, which you were staggering anyway. Best balance of rigour and cost in an HR setting. It earns its place because it's the only common method that separates the tool from the weather. Both groups sit inside the same market, the same pay review, the same freeze, so a difference in one and not the other has a real claim to be the software. Weak because the held-back group knows it's held back, and behaviour drifts. A manager who wants the tool will lobby hard to be in wave one. The specific failure: your control group is the region that was struggling, which is exactly why it was scheduled second. Reads in one to two quarters.

Time-and-motion sampling of the task itself

Watch a small sample of the actual work, before and after, and time it. Sit beside two recruiters for a morning with a stopwatch. Then repeat it once the tool has bedded in. Not the system log, the real work, including the second look at a CV that the log records as four seconds. Cuts through the noise in system logs. It earns its place because it sees the parts of the job no system records, which is usually where the tool either helps or quietly doesn't. Weak because samples are small, and the act of watching changes how people work. Someone being observed skips the tea round and the third pass at the wording. The specific failure: you time the new process at its best, while the baseline it's measured against contained every ordinary interruption, and you report the difference as the tool. Reads in weeks.

Quality outcomes measured downstream

Track first-year attrition of hires, error rates, case reopen rates. The only one that proves the tool didn't break the thing it sped up. A screening assistant that halves shortlisting time while quietly narrowing the pool will look excellent for a year, then arrive as a retention problem nobody traces back to the software. It earns its place for that reason alone: it catches damage rather than speed. Weak because the read time is a year or more, longer than most renewals. You'll be asked to sign long before that cohort has been in post long enough to tell you anything. The practical move is to start the clock now and say so in the paper, so the read exists for the next renewal. The specific failure: the metric was never defined at rollout, so the cohort has nothing to be compared against. Reads in twelve months plus.

Staff-reported time saved

Survey the people doing the work, before and after. Cheap and often the most honest signal, particularly if you ask about last week rather than in general, and ask what the person did with the time rather than how much they saved. It earns its place because it's the only method that surfaces the workaround. When a coordinator tells you she re-reads every AI-drafted offer letter by hand before it goes out, you've learned something no dashboard will show you. Weak because self-reports inflate, and the people who hate the tool also tend to overstate the time it costs them. Enthusiasts inflate the other way. Neither group is lying. Both are reporting a feeling rather than a duration. The specific failure: you run the survey the week after the town hall where the rollout was called a success, and every answer comes back pre-shaped. Reads in a quarter.

The vendor's own reported benchmark

What the platform tells you. Useful as a directional signal, not as a basis for a board paper. It earns a place because it costs nothing and moves daily, so it shows a collapse in usage long before anyone on your team reports one. If the deflection rate falls off a cliff in week six, you want to know that on the day. Weak because the vendor defines the metric, chooses the cohort, and reports the average. Deflection counts the conversation the assistant ended, not the employee who gave up and emailed a person. Hours saved is usually a rate the vendor picked, multiplied by a volume the vendor counted. Use it to spot movement, never to size a saving. The specific failure: the figure goes into your board paper carrying the vendor's definition, and somebody asks what a deflected ticket actually is. Reads whenever you log in.

The Decision Table

Situation Scale Setup Primary Pain Recommended Starting Point
No AI tool in place yet Under 1,000 headcount Manual or single-system Looking for an efficiency story Time-and-motion sampling on the current process before any procurement
Tool in place, renewal under 12 months away 1,000 to 5,000 AI plus legacy Cannot prove the value internally Held-back control group in the next wave, or peer-team comparison now
Tool in place, no baseline, renewal in 60 days 1,000 to 5,000 AI plus legacy Renewal meeting with no slide After-the-fact reconstruction plus staff-reported time saved
Procurement for compliance, not efficiency Any Regulated industry Audit pressure, not efficiency Step out of the ROI frame, route to legal and audit
Tool in place, large spend, board reporting 5,000 plus AI integrated Career risk on a soft number Quality outcomes measured downstream, accepted as a long read
Tool being considered for one HR function only Under 1,000 One function, mostly manual Function head is the champion Pilot-first measurement in that function only
Renewal of a tool that the team has stopped using Any AI plus legacy Low adoption, high renewal cost Diagnose non-use before measuring ROI, the cost is in the spend, not the missing number
Tool being justified to a non-HR buyer (CFO, CEO) 1,000 to 5,000 AI plus legacy Cross-functional scepticism Before-and-after on a single metric, paired with a quality outcome the buyer already cares about

The Cost of Getting This Wrong

A wrong number doesn't cost the budget. It costs the next conversation. Once the CFO has seen a soft ROI on an HR tool, every HR business case for the next three cycles is read with one more layer of suspicion. The second-order cost isn't a line item. It's the trust account you draw from every time you ask for headcount, for a new tool, for a quiet hire. And the tool vendors know this, which is part of why their decks are so confident and so vague at the same time.

There's a version of this that hurts more, because it looks like a win. The number holds. The renewal goes through. Now the claim is load-bearing. It gets repeated in a board pack by someone who wasn't in the original room, it becomes the reason a second module is bought, and by the time anyone tests it the person who built the figure has moved on. Wrong numbers that fail get corrected. Wrong numbers that succeed get built on.

There's also a quieter cost inside the team. Recruiters who know the chatbot is misrouting tickets and have been told to celebrate the deflection rate stop volunteering bad news. That's the moment an HR function stops learning. The signal doesn't come back later, either. A coordinator who raised the misrouting twice and watched it go nowhere won't raise the next thing, which might be an offer letter template that's been wrong for a fortnight.

And one person carries a cost nobody counts. Whoever sponsored the tool now owns the story rather than the question. Every honest read of the data becomes an admission, so honest reads stop being commissioned. The sponsor ends up defending software they privately suspect. That isn't a character flaw. It's what happens when you tie a person's judgement to a number nobody measured independently. So the question to sit with, going into the renewal, isn't "did the tool work?" It's "what did I stop measuring because the tool was easier to point at?"

When You Are Ready to Go Further

If you've read this far and the decision is still a coin flip, the gap is usually not motivation. It's method. A second opinion from someone who has watched eight of these renewals go wrong and two go right is worth the time.

HROpsLab is an independent review publication. We don't sell software, we don't run implementations, and we don't take referral fees from the vendors we write about. Our comparison work is built for the reader who has to defend a decision in a room where the vendor isn't allowed to speak for them. If that's the situation you're in, we are worth a conversation.


Frequently Asked Questions

What should I measure first if I have nothing yet?

Time-and-motion sampling on the current process, before any procurement, so you've a number to compare any future claim against. Pick the one process people complain about. Sit with two of them for a morning and record how long the steps really take rather than how long the system says they took. The work to measure a process you already run is small. The work to measure a process you've already changed is much larger, because the old version exists nowhere except in memory, and memory bends towards whatever story the room is telling.

How do I build a baseline after the tool is already live?

Reconstruct it from old tickets, archived requisitions, prior quarterly reports, and a peer team that has not been switched on yet. Pull a fixed window from before rollout, use the definitions that were in force back then rather than the ones the tool introduced, and write down what you couldn't recover. It won't be as clean as a pre-rollout capture, but it's defensible, and defensible beats absent every time. What saves you in the meeting is naming your own gaps first. A reconstruction that admits its edges reads as careful work.

Is staff-reported time saved real money?

Often not, because the time gets redeployed into more sourcing calls, more candidate touchpoints, more review work. Ask where the saved hours went before you convert them into a dollar figure on a slide. The follow-up is the question that decides it: what did you stop doing? If nobody can answer, the hours were absorbed rather than saved, and the saving exists only in your business case. Time turns into money when someone with authority decides not to backfill a role, or lets a contract lapse. Until then, you have capacity.

How long should a pilot run?

A full cycle of the work it touches, which is usually a quarter for recruiting and a quarter for HR service. Shorter than that and you're reading ramp-up, not steady state, and the number will flatter the tool. Make sure the window contains at least one genuinely busy period, because a service desk measured in a quiet month is a service desk you haven't measured. The other trap is stopping on a high. End the pilot the week after a strong month and that's the figure that goes in the deck.

What do I do when the vendor benchmark disagrees with mine?

Put both on the same slide, name the difference, and say which one is yours. Before that, find out what sits underneath theirs: how many customers are in the average, and what counts as a resolved case. Most of the gap is usually definition rather than performance, and being able to demonstrate that turns an apparent contradiction into a footnote. The disagreement is the conversation, and a CFO who sees you handle it openly is a CFO who trusts the next number you bring. Leaving it out is worse than showing it.

How do I report a small gain honestly?

Frame it as a small, defensible gain on a single process, paired with a quality outcome you're willing to wait for. Show the method before the result, so the room judges your work rather than your conclusion. Say what you couldn't measure. Say what would change your mind. A number that arrives with its own limits attached is harder to attack than one that arrives alone, and whoever brings it keeps the standing to bring the next one. "Small and real" beats "large and soft" at every budget review the author has ever watched.

We help HR operations leaders make defensible decisions on the tools they already run.

Share on X Share on LinkedIn

What to do next?

Explore More Articles

Dig deeper into HR Ops strategy, tools, and workflows built for real teams.

Browse the blog →
Join the HROpsLab Community

Connect with People Ops practitioners sharing real workflows, tools, and challenges.

Join now →