TL;DR
- The number that matters is not usage. It is how many questions stopped reaching a person, and most teams never baseline it before go-live, which makes the after number meaningless.
- Week one is about wrong answers, not volume. Find them deliberately while the audience is small enough that a wrong answer is embarrassing rather than damaging.
- Three things break in a predictable order: stale documents, missing escalation, and questions nobody thought to write down.
- Set the bar before you start. "It is working" with no number attached becomes an argument in month three when somebody asks what it cost.
- The decision rule at day 30: if deflection is below ten per cent and the content is current, the tool is wrong. If the content is not current, the tool was never the problem.
- The outcome: a defensible answer to "is this worth keeping", backed by two numbers you can show somebody.
The Launch That Nobody Measured
An HR team switches on an AI agent trained on their handbook. It answers questions. People use it. Three months later the renewal arrives and somebody asks whether it is working.
The honest answer is that nobody knows. There is a dashboard showing four hundred conversations, which sounds like a lot until you realise nobody counted how many questions HR was fielding before, so there is nothing to compare it against. The four hundred might be four hundred questions that would have reached a person, or four hundred people poking a new toy, or both.
That is the failure this piece is written to prevent, and it is almost entirely avoidable with one afternoon of counting before anybody switches anything on.
Best tools for Device Management
The thing to understand about the first thirty days is that it is not a performance test. The tool will answer questions from day one. What the thirty days actually tell you is whether your documentation is good enough, whether your escalation path works, and whether the volume you were hoping to remove was ever the volume you had. Those are three different findings and only one of them is about software.
Before You Switch Anything On
Decide who is actually running this
Name one person before anything else. Not a committee and not a role. The first month has a daily task in week one and a weekly one after that, and a deployment with no named owner reliably produces four hundred conversations and no findings, because logging a wrong answer is nobody's job in particular.
That person needs twenty minutes a day for a week and about two hours a week after. If they do not have it, the honest move is to delay the launch rather than start and let the measurement lapse, because an unmeasured first month cannot be recovered retrospectively.
Baseline the volume you are trying to remove
One month of requests, counted. Not estimated. Mark each as answerable-from-existing-documents or genuinely needing a person. That first number is the ceiling on what any deflection tool can remove, and if it is small, no amount of configuration will make the project look good.
If you did this before choosing a tool, you already have it. If you did not, do it now and accept that your baseline and your launch overlap slightly. An imperfect baseline beats none.
Write down what "working" means, with a number
Not "fewer questions". Something you can check: "by day thirty, at least a third of questions that used to reach HR are answered without us". Pick the number before you have a result, because a number chosen afterwards is a justification rather than a target.
Check what the tool counts as a conversation
Worth ten minutes before launch because it determines whether your headline number means anything. Vendors count differently: some log every message, some every session, some every unique person per day. A tool counting messages will report a number three or four times larger than one counting sessions, for identical usage.
Ask the vendor directly, write the answer down, and use the same definition all the way through. The specific trap is comparing a month-one figure against a month-six figure after the vendor has changed how the dashboard aggregates, which happens more often than you would expect and is never announced.
Pick the five questions you will test every week
The five HR answered most often last month. Ask them on day one, day seven, day fourteen and day thirty, and read the answers against the actual policy documents each time. This is the single most useful recurring check in the whole process and it takes ten minutes.
Agree who owns a wrong answer
Before launch, not after. When the agent tells somebody the wrong notice period, who corrects the document, who tells the employee, and who decides whether it was serious. Name those people. The first wrong answer is not a crisis if somebody owns it and is a mess if nobody does.
Decide what it will not be asked
Grievances, anything involving a named colleague, anything about an individual's pay or performance. Those should route to a person immediately, by design, and the agent should say so plainly rather than attempting an answer.
The Thirty Day Schedule, on One Page
The weeks below are described in detail further down. This is the version to put in a calendar, because the main reason first months drift is that nobody allocated the time rather than that anybody disagreed with the approach.
| When | Task | Time | Who |
|---|---|---|---|
| Before day 1 | Count one month of requests, split deflectable versus not | Half a day | Whoever reads the inbox |
| Before day 1 | Write the definition of working, with a number | 15 min | HR lead |
| Before day 1 | Pick the five test questions, test the exclusion path | 30 min | HR lead |
| Day 1 | Announce once, saying what it can and cannot do, and that it is AI | 15 min | HR lead |
| Days 1 to 7 | Hunt wrong answers, log each with its source document | 20 min daily | Named owner |
| Day 7 | Read the wrong-answer log, sort by question frequency | 30 min | Named owner |
| Days 8 to 21 | Fix documents by frequency, re-ask after each fix | 2 hours weekly | Document owner |
| Days 8 to 21 | Check where requests are still arriving | 10 min weekly | Named owner |
| Day 21 | Capture the undocumented-knowledge list somewhere permanent | 1 hour | HR lead |
| Day 30 | Deflection rate, wrong-answer cause split, the one question | 1 hour | HR lead |
That totals roughly two days of effort spread across a month, most of it in the first week. The part teams skip is the half day before day one, and it is the part that makes everything after it interpretable.
One scheduling note. Do not start this in a month that contains your busiest HR period. A first month that overlaps with annual review season or open enrolment will produce a volume spike that makes the baseline useless and the deflection figure flattering for the wrong reason.
Week One: Go Looking for Wrong Answers
The instinct in week one is to watch adoption. Resist it. Adoption in week one is curiosity and tells you nothing. What week one is for is finding the wrong answers while the audience is still small.
Ask it the five questions, and read the answers against the documents. Not against your memory. Memory is where stale policy lives.
Ask it something that is deliberately not in the documentation. The behaviour you want is an honest handoff with the conversation attached. A confident invention is disqualifying, and better discovered now than by two hundred people.
Ask it something sensitive. "I want to raise a grievance about my manager." It should route to a person without attempting advice. If it answers with policy text, fix that configuration before anything else.
Have somebody outside HR ask five questions in their own words. Staff do not phrase things the way the handbook does, and the gap between "what is the carry-over limit" and "can I keep my leave" is where most early failures sit.
Pick that person deliberately. Somebody who joined recently is ideal, because they have no accumulated knowledge of where things live and no habit of asking a particular colleague. A long-tenured employee will unconsciously phrase questions the way the handbook does, having read it years ago, and will therefore fail to find the gaps that matter.
Log every wrong answer with the document behind it. Two columns: what it said, which document it came from. By day seven that list is your content review backlog, prioritised by what people actually ask.
What a good week one looks like
Between three and eight wrong answers found, most of them traceable to a document that is out of date rather than to the tool misreading a current one. That ratio is the useful signal. If nearly every wrong answer comes from a correct document, you have a tool problem. If nearly all come from stale documents, you have a content problem the tool just made visible.
| Week one finding | What it actually means | What to do |
|---|---|---|
| Wrong answer, document was stale | Content problem, newly visible | Fix the document, keep the tool |
| Wrong answer, document was current | Tool is misreading your content | Raise it with the vendor, test strict mode |
| Confident answer to something undocumented | Configuration problem, serious | Fix before wider rollout |
| Sensitive question answered rather than routed | Configuration problem, urgent | Fix immediately |
| No wrong answers at all | Nobody is really using it yet | Check where staff are still asking |
Weeks Two and Three: Fix the Documents, Not the Tool
By week two the wrong-answer list is the work. Most of it will be documents rather than software, which is the finding that surprises teams and is also the most valuable thing the first month produces.
Work the list by frequency, not by severity. The question asked thirty times a month on a document nobody has reviewed since 2024 matters more than the edge case asked once.
Re-ask after each fix. A corrected document does not always produce a corrected answer immediately, because some tools cache or need a re-crawl. Confirm rather than assume.
Watch where people are still asking. If requests are still arriving in the shared inbox at the old rate, the problem is not answer quality, it is that nobody knows the agent exists or it is in the wrong place. A tool living on a web portal in a company that lives in Slack will be used by HR and nobody else.
Start counting deflection properly. Questions the agent answered without a handoff, as a share of total questions reaching HR by any route. The denominator matters: if you count only conversations the tool saw, you will report a flattering number that means nothing.
Resist adding features. The temptation in week two is to configure more: categories, workflows, extra sources. Everything added is something to maintain, and you are thirty days from knowing whether the basic thing works.
There is one exception worth making. If week one showed the agent confidently answering something undocumented, fix that configuration immediately rather than waiting, because every day it runs is a day it may tell somebody something invented. Everything else on the wish list can wait until day thirty.
Tell people what changed. A short note in week three saying which questions now get answered instantly, naming two or three specifically, does more for adoption than the launch announcement did. Staff who tried it in week one and got a stale answer will not come back unprompted, and that cohort is the hardest to recover later.
Week Four: Read the Numbers Honestly
At day thirty you need two numbers and one judgement.
Deflection rate. Questions answered without a person, over total questions reaching HR by any route. Compare against the baseline ceiling you calculated before launch, not against zero.
Wrong-answer rate, and its cause. How many of the answers you spot-checked were wrong, and whether the cause was stale content or tool error. This is the number that decides whether to keep the tool or fix the handbook.
The judgement: did the work actually move? Ask the person who used to field the questions whether their week feels different. It is subjective and it is the single best indicator, because a deflection rate that does not change how anybody's week feels is measuring the wrong thing.
Ask it as an open question rather than a yes or no. "What is different about your week" produces something usable; "is your week better" produces agreement, because the person answering knows what you are hoping to hear and knows the tool was your idea. If the answer is specific, such as no longer being interrupted during the Monday payroll run, that is a real result you can report. If it is general approval with no example attached, treat the deflection number with more suspicion than you otherwise would.
Reading the result
| Deflection at day 30 | Content current? | Honest conclusion |
|---|---|---|
| Above 30 per cent | Yes | Working. Keep it, keep reviewing the content |
| 10 to 30 per cent | Yes | Working, under-adopted. Fix placement before blaming the tool |
| Below 10 per cent | Yes | The tool is a poor fit, or the volume was never deflectable |
| Any figure | No | The content was the problem. Fix it before judging the tool |
| Above 30 per cent | Yes, but nobody's week changed | You deflected questions nobody minded answering |
That last row is the uncomfortable one and it is more common than people expect. Removing forty easy questions a month can leave the actual burden untouched, because the burden was six difficult cases rather than forty easy ones. That is not a failure of the tool, it is a sign the original diagnosis was wrong, and it is better known at day thirty than at renewal.
The Three Things That Break, in Order
Stale documents, week one. Every deflection tool answers from content somebody has to keep current. The first month of any deployment is, in practice, a content audit conducted by your employees on your behalf.
This is worth reframing as a benefit rather than a setback, because it is the most reliable return the first month produces. A handbook review that nobody has prioritised for two years happens in three weeks when employees are actively finding the gaps and the gaps have names attached. The tool is doing something genuinely useful even in the month where its own numbers look poor, and that is worth saying out loud to whoever approved the spend.
Missing escalation, week two. The handoff works in testing and fails in reality, usually because the person it routes to was not told, or the conversation context does not travel with it, so the employee repeats themselves and concludes the whole thing is useless.
The detail that causes this is almost always mundane. The handoff lands in a channel nobody watches, or it arrives as a notification that looks like every other notification, or it goes to a shared address rather than a person. Test it by having somebody genuinely escalate and timing how long before a human responds. If that number is longer than the old inbox, the deployment has made things worse for the cases that matter most, which is the opposite of the intended outcome and entirely invisible in a deflection metric.
Questions nobody wrote down, week three. The agent handles everything in the handbook and then meets the real volume: the things everybody knows and nobody documented. Which manager approves what. How the expenses thing actually works now. These are not failures of the tool, they are a map of your undocumented knowledge, and that map is arguably worth more than the deflection.
Five Questions Teams Ask in the First Month
"Should we announce it or let people find it?" Announce it, once, with what it can and cannot do. Quiet launches produce low adoption and then a false conclusion that the tool does not work.
"What if it gives bad advice?" It will, in week one, which is exactly why week one is for finding that deliberately with a small audience. The plan is a named owner for corrections, not an expectation of perfection.
"Do we tell people it is AI?" Yes. It is the straightforward thing to do and it also sets expectations, which reduces the frustration of a handoff. Staff are considerably more patient with a bot that admits what it is.
"How do we stop it answering sensitive things?" Configuration, tested before launch. Ask it a grievance question and confirm it routes rather than advises.
"When do we know it is not working?" Day thirty, against a number you wrote down beforehand, with the content question answered first. Without that, "not working" is a feeling and the conversation goes nowhere.
What the Dashboard Will Not Tell You
Every tool reports what it saw. That is the whole limitation.
It cannot count the questions asked in a corridor, in a direct message to a colleague who is not in HR, or in a reply to an unrelated thread. In most companies those are a real share of volume, and they are invisible to the tool precisely because they never reached it. A dashboard reporting four hundred conversations in an organisation generating nine hundred questions is not lying, it is partial, and the gap is the thing you most need to know.
It also cannot tell you whether an answer was correct. It reports that it answered, not that the policy it quoted was current. Accuracy is a spot-check somebody performs, not a metric a tool produces about itself.
And it cannot tell you whether the work moved. Deflection is a proxy. The real question is whether the person who used to answer those questions now spends that time on something better, and the only way to find out is to ask them.
The diagnostic at day thirty is one question. Ask the person whose inbox this was supposed to fix whether their week is different. If yes, the numbers are confirming something real. If no, the numbers are measuring activity rather than outcome, and it is worth asking what the actual burden was before renewing anything.
The Decision Table
| Situation | Scale | Setup | Primary Pain | Recommended Starting Point |
|---|---|---|---|---|
| No baseline taken, already live | Any | Any | Nothing to compare against | Count one month now, accept the overlap |
| Wrong answers, documents stale | Any | Any | Content, not software | Fix documents by question frequency |
| Wrong answers, documents current | Any | Any | Tool misreads your content | Vendor conversation, test strict mode |
| Low usage, answers are good | 100 plus | Portal-based | Nobody knows it exists | Move it to where staff already are |
| Good deflection, week feels the same | Any | Any | Wrong problem diagnosed | Re-examine what the real burden was |
| Sensitive questions being answered | Any | Any | Configuration, urgent | Fix before any wider rollout |
| Handoffs arriving without context | Any | Any | Escalation path incomplete | Fix the handoff before adding sources |
Rows two and three look identical from the outside and have opposite remedies, which is why the wrong-answer log needs the document name next to it. Without that column you cannot tell a content problem from a tool problem, and teams routinely replace a working tool because of a stale handbook.
What Happens After Day Thirty
The first month gets all the attention and then the process stops, which is why deployments that looked fine in month one quietly decay by month six. Three things need a home in somebody's calendar.
The five test questions, monthly. Ten minutes. The same five questions, read against the current documents. This is the cheapest possible guard against silent decay, and decay is the normal state: policies change, documents get edited by people who do not know a bot reads them, and nobody notices until an employee acts on something wrong.
The deflection rate, quarterly. Not monthly, because the number is noisy at that interval and you will over-react to seasonality. Quarterly it tells you whether the thing is still earning its cost, and it gives you a figure for the renewal conversation that was not invented the week before.
The wrong-answer log, kept open permanently. It is not a launch artefact. Every wrong answer found in month eight has the same two columns and the same value, and the log is the only record that distinguishes a content problem from a tool problem over time. Teams that close it at day thirty lose the ability to tell those apart later, which is exactly when the distinction becomes expensive.
What changes in month two and three. Volume usually dips after the launch curiosity passes, then settles higher than week one. Expect that shape and do not read the dip as failure. What should also appear by month three is a second category of question the agent cannot handle, different from week one's: not undocumented knowledge, but genuine judgement calls that were always going to need a person. Those are the floor. If the remaining volume is all judgement calls, the tool has done its job and further configuration will not help.
When to re-run the whole thirty days. After a handbook rewrite, after a change of HRIS, or after a reorganisation that changes who approves what. Each of those invalidates a share of the source content at once, and a tool answering confidently from superseded documents is worse than no tool, because people trust it now.
Where Teams Get the First Month Wrong
Measuring usage instead of deflection. Conversations is a vanity number. The question is how many questions stopped reaching a person, and that needs a denominator the dashboard does not have.
No baseline. Without a before number, the after number cannot be interpreted, and the renewal conversation becomes an argument about impressions.
Launching quietly to avoid embarrassment. Produces low adoption, which produces a false negative, which kills a tool that was working.
Treating wrong answers as a tool failure by default. Most early wrong answers are correct readings of out-of-date documents. The log needs the document name or you will reach the wrong conclusion.
Adding sources and features in week two. Every addition is maintenance, and you do not yet know whether the simple version works.
Forgetting to tell the person on the other end of the handoff. The escalation path is the part that fails in reality and works in testing, and the failure looks like a bad bot rather than a missing briefing.
Judging it before the content is current. If the handbook has not been reviewed, the first month measures your documentation, not the software. That is useful, but it is not an evaluation of the tool and should not be reported as one.
What to Establish Before Day One
The baseline count. One month of requests, split into deflectable and not.
The definition of working, with a number. Written down before launch.
The five test questions. The ones HR answered most often last month.
The named owner of a wrong answer. A person, not a role.
The escalation path, tested. Including whether the receiving person was told and whether context travels.
The exclusion list. What it will route rather than answer, confirmed by asking it.
Where staff will actually find it. Observed from where last month's requests arrived, not from where you would prefer.
None of those seven needs the vendor. They are answerable inside your own organisation in a day, and they are the difference between a deployment you can defend at renewal and four hundred conversations nobody can interpret.
What Getting the First Month Wrong Costs
The first cost is a tool cancelled for the wrong reason. A deflection rate of eight per cent against a stale handbook says nothing about the software, and a team that reads it as a product failure replaces a working tool and repeats the experience with the next one.
The second is the opposite error and it is more expensive over time. A flattering dashboard, no baseline, nobody's week actually different, and a subscription renewing annually on the strength of a number that was never connected to the work. This is the quieter failure and it survives for years.
The third costs trust rather than money. A sensitive question answered with policy text instead of routed to a person is not a configuration note to the employee who asked it. That damage is slow to repair and the test that prevents it takes two minutes before launch.
The fourth is the undocumented knowledge that never gets written down. Week three produces a list of things everybody knows and nobody has recorded, which is genuinely valuable, and in most deployments that list is read once and lost. Capture it somewhere permanent, because it outlives whichever tool you are currently evaluating.
There is a fifth cost that only lands at renewal and is worth naming because it is avoidable with one afternoon. Without a baseline and a written definition of working, the renewal decision gets made on whoever argues most confidently rather than on evidence, and that conversation tends to happen under time pressure with a finance deadline attached. The teams who sail through it are not the ones whose tool performed best, they are the ones who can produce two numbers and say where they came from.
When You Are Ready to Call It
At day thirty, answer three things in order. Is the content current? If not, stop, fix it, and judge the tool in another month, because right now you are evaluating your handbook. If yes, what is the deflection rate against the baseline ceiling? And does the person whose inbox this was meant to fix say their week is different?
Those three answers produce a defensible decision rather than an impression. Keep it, move it somewhere more visible, fix the content first, or accept that the volume you wanted to remove was never the volume you had.
And whichever way it goes, keep the wrong-answer log and the undocumented-knowledge list. Both outlast the tool, both are useful to the next person, and neither can be reconstructed later from a dashboard.
One last thing worth saying plainly, because it is the finding teams least expect. A first month that ends with a poor deflection rate and a handbook that is finally current is not a failed project. It is a content audit you were never going to schedule, completed by your own employees, with the gaps named and prioritised by what people actually ask. Judge the software in the second month, once that work is done, and judge it against the baseline you wrote down rather than against the hope you started with.
For a side-by-side comparison of the tools themselves, including published pricing and where each one loses, see our [HR chatbot software comparison](/best-hr-chatbot/). If you are earlier than that and still deciding whether you need deflection or case tracking, start with [HR help desk software](/hr-help-desk-software/).
Frequently Asked Questions
How long before an HR chatbot shows results?
Answers start on day one, but a result you can defend takes about thirty days, and what that month actually measures is usually your documentation rather than the software. The sequence is predictable: week one surfaces wrong answers, weeks two and three are spent fixing the documents behind them, and week four is the first point at which a deflection number means anything. The teams who report a result in the first week are reporting conversation counts, which is activity rather than outcome, and that number cannot be interpreted without a baseline taken before launch.
What should we measure in the first month?
Two numbers and one judgement. Deflection rate, meaning questions answered without a person as a share of total questions reaching HR by any route, not just the ones the tool saw. Wrong-answer rate with its cause recorded, specifically whether each wrong answer came from a stale document or from the tool misreading a current one, because those have opposite remedies. Then ask the person whose inbox this was meant to fix whether their week feels different. That last one is subjective and it is the best single indicator, because deflection that changes nobody's week is measuring the wrong thing.
What if the chatbot gives a wrong answer?
Expect several in week one, which is exactly why week one should be run with a small audience and a deliberate hunt for them. Log each one with the document it came from, because that column is what tells you whether you have a content problem or a tool problem, and most early wrong answers are correct readings of out-of-date policy. Agree before launch who corrects the document, who tells the employee, and who judges severity. A wrong answer with a named owner is a routine fix, and the same wrong answer with nobody responsible becomes an argument about whether the tool should be switched off.
Should we announce the chatbot or launch quietly?
Announce it once, saying plainly what it can and cannot do, and that it is AI. Quiet launches produce low adoption, low adoption produces a false negative, and the false negative kills a tool that was working fine. Being upfront that it is AI also sets expectations, which makes people considerably more patient when a question gets handed to a person. The one thing worth doing before the announcement is confirming it routes sensitive questions rather than answering them, because the first impression you cannot undo is a grievance met with policy text.
How do we stop it answering sensitive questions?
Configure an exclusion path before launch and then test it by asking, in plain language, something like wanting to raise a grievance about a manager. The correct behaviour is an immediate route to a named person with no attempt at advice. If it responds with policy text instead, treat that as the most urgent item in the deployment, ahead of answer quality or adoption. Also test what a colleague without permission can see when they search for a sensitive case, because access restriction and answer routing are separate settings and teams commonly configure one and assume the other.
What is a good deflection rate?
Above thirty per cent at day thirty is a solid result, between ten and thirty suggests it works but is under-adopted, and below ten per cent means either the tool is a poor fit or the volume was never deflectable in the first place. All of those readings depend on the content being current, so answer that question first: a low rate against a stale handbook tells you nothing about the software. And compare against the baseline ceiling you calculated before launch, which is the share of requests that had a written answer at all, rather than against zero or against a figure from somebody else's case study.
Why did deflection go up but nothing feel different?
Because the questions removed were not the ones carrying the weight. Forty easy repeat questions can disappear while the real burden stays untouched, if that burden was six difficult cases rather than forty simple ones. This is a diagnosis error rather than a tool failure, and it is worth catching at day thirty rather than at renewal. The useful response is to go back to the original request log and ask which items actually consumed time, because the answer frequently points at case management or at a process problem rather than at anything a chatbot can address.
What do we do with the questions nobody documented?
Write them down, because that list is often the most valuable thing the first month produces. By week three the agent will have handled everything in the handbook and started meeting the real volume: which manager approves what, how a process actually works now rather than how it was written, the things everybody knows and nobody recorded. Capture those somewhere permanent rather than in the tool's own logs, because the list outlives whichever product you are currently evaluating and it cannot be reconstructed from a dashboard afterwards.