TL;DR
- Every HR answer bot has a set of questions it should decline, and the useful question when buying one is whether it declines them or answers them anyway.
- If your HR questions are all procedural and none of them depend on individual circumstances, you do not need this. Most companies are not in that position.
- Four categories need a person regardless of how good the tool is: anything depending on individual circumstances, anything jurisdiction-specific without a local document, anything requiring action rather than an answer, and anything someone discloses rather than asks.
- Answering rather than actioning is a real limit and not a flaw. A tool that explains a process cannot complete it, and buying one expecting the second is the most common disappointment in this category.
- A vendor that states its limits plainly is a safer purchase than one that does not, because the limits exist either way and only one of them has thought about the handover.
- Configure the refusals before launch. The default behaviour of most tools is to answer, and the default is wrong for the questions that matter.
The Question It Should Have Refused
An employee asked an HR answer bot whether taking a particular kind of leave would affect an upcoming performance review. Sensible question, nervously asked, in the employee's second language, late in the evening.
The tool answered it. The answer was assembled from two documents, one about the leave process and one about the review cycle, and it was reasonable, fluent and confident. It was also a judgement rather than a fact, and it was a judgement nobody at the company had made. No document said anything about the interaction between those two things, so the tool had done what these tools do when the material is thin: it had produced the most plausible-sounding synthesis available.
The employee acted on it. Not dramatically, but they made a decision about timing, and they did not raise it with anybody because they had already been told. The problem was not that the answer was wrong, and in fact nobody ever established whether it was. The problem was that a question requiring somebody to take responsibility for an answer was answered by something that cannot take responsibility for anything. A refusal would have been a better outcome than a correct answer, because a refusal routes the question to a person who can own it.
Best tools for AI HR Tools
So the capability worth evaluating is not what the tool can answer. It is what it declines to.
When You Don't Need to Think About Refusals
When the manual way is genuinely fine
Your HR questions are genuinely all procedural: where to find things, how to submit things, what the dates are. If that is true, the refusal question is theoretical, and a tool answering everything is fine because nothing in the set is consequential. Be honest about whether it is true, though, because most question sets contain a small sensitive minority that nobody has looked for.
When friction starts appearing
The signal is a tool giving answers that are technically supported by your documents and that nobody at the company would have given in those words. Often spotted by accident, in a transcript review or because an employee quotes it back. The useful response is to stop treating it as an accuracy problem and start treating it as a scope problem.
When it becomes a liability
The point where somebody has relied on a judgement the tool made on your behalf. The exposure is specifically that the company now has a position it never decided to take, delivered in writing, repeatedly, to anybody who asked a similar question. This is worse than an individual wrong answer because it is systematic.
The edge case that forces it
A restructure, a change programme, or anything that generates anxious questions about individual consequences. These question sets are almost entirely in the refuse-and-route category, and they arrive in volume at precisely the moment a self-service tool looks most attractive as a way of absorbing load.
Five Questions People Ask First
"Isn't a refusal just the tool failing?" No, and this framing is why so many deployments are configured badly. A refusal that names the right person is a successful outcome, and it should be measured as one. If your reporting counts a routed question as a deflection failure, your metrics are pushing the configuration in exactly the wrong direction.
"Can't we just train it on more documents?" For the coverage gaps, yes, and that genuinely helps. For the four categories below it does not, because those questions are not answerable from documents at all. Adding material makes the tool more confident in the areas where confidence is the problem.
"How do we know what it is refusing?" Ask for reporting on routed conversations as a specific category, not buried in an unresolved bucket. If a tool cannot tell you what it declined and why, you cannot tune the behaviour or spot a refusal rule that is firing too broadly.
"Will people be annoyed by refusals?" Much less than by wrong answers, provided the refusal is useful. "I cannot answer that because it depends on your circumstances, here is who can" is a good employee experience. "I do not understand the question" is a bad one, and the difference is entirely in the wording and the routing rather than in the decision to decline.
"Does it have to be configured, or does it come that way?" It has to be configured, almost always. The default behaviour across this category is to answer, because answering demonstrates capability and demonstrations sell software. Treat the refusal set as implementation work you own.
The Four Categories That Need a Person
Anything that depends on individual circumstances
The largest category and the one that generated the opening example. Questions whose correct answer varies by the asker's contract, tenure, location, role or situation. A document-trained tool sees a question it has material adjacent to and produces a general answer with the specificity stripped out, which reads as applying to the asker because they asked it.
The configuration you want is a refusal keyed to the topic rather than to the phrasing, because people ask these in endlessly varied ways and a phrase-based rule will miss most of them.
Anything jurisdiction-specific with no local document
Where the answer genuinely differs by country and you have not written the local version. The tool will answer from the general document, unqualified, because it has no way to know that a local variation exists but is missing. This is the single most common source of confidently wrong answers in a multi-country deployment.
The fix is partly configuration and mostly content: a short written section for each divergent topic saying that the specifics depend on where somebody is employed and naming who handles it. That gives the tool something correct to retrieve instead of something plausible to synthesise.
Anything requiring action rather than explanation
Changing a benefit election, booking leave, correcting a payroll detail, updating a record. An answer bot explains how; it does not do. This is the limit most often misunderstood at purchase, because a demo showing a tool explaining a process looks very like a tool completing one.
Nothing to configure here beyond honesty: the tool should say what it can and cannot do and hand over cleanly. If your requirement is genuinely that the request gets processed, you are looking for a workflow platform and this is the wrong purchase.
Anything disclosed rather than asked
Somebody telling you something, rather than asking. A health circumstance, a problem with a colleague, a safety concern, distress. These frequently arrive phrased as questions, which is exactly why they need handling as a category rather than by intent detection on the question mark.
This is the one place where a conservative rule is clearly right: any conversation touching these topics should route to a named person immediately, with no attempt at an answer, and the tool should say plainly that a person will pick it up. Over-routing here costs very little and under-routing costs a great deal.
| Category | Why a tool cannot handle it | What to configure | Cost of getting it wrong |
|---|---|---|---|
| Depends on individual circumstances | The general answer reads as applying to the asker | Topic-level refusal, not phrase-level | A judgement the company never made, delivered in writing |
| Jurisdiction-specific, no local document | It answers from the general document, unqualified | A written routing section per divergent topic | Confidently wrong answers, systematically |
| Requires action, not explanation | It explains how. It cannot do | Honest statement of scope plus a clean handover | A purchase that disappoints regardless of answer quality |
| Disclosed rather than asked | Needs a person from the first message | Conservative routing on the whole category | Over-routing costs little. Under-routing costs a lot |
Writing the Refusals
A refusal is content, and good refusal content has four parts.
Say that you cannot answer, and why, in one sentence. The why matters: "because this depends on your individual circumstances" tells the employee something true and useful, and distinguishes the refusal from a failure to understand.
Name the route. A person or a monitored inbox, not "contact HR". The whole value of the refusal is the handover, and a handover to a generic destination is where routed questions go to wait.
Set the expectation. When they will hear back, in their working days. A refusal with no timing attached converts a fast interaction into an indefinite one, which feels worse than the wait actually is.
Carry the context across if the tool can. The best behaviour is for the routed conversation to arrive with the original question attached, so the employee does not retype their situation. Ask whether the tool does this; many do not, and it is the difference between a handover and a dead end.
Write these once per category rather than per question. Four well-written refusals cover most of the surface, and they are reusable across languages, which matters because a refusal is the one piece of content you most want to be correct in every language you support.
What Happens to the Routed Question
Everything above treats routing as the destination. For the person receiving routed questions it is the start, and the arrangement usually fails on their side rather than on the tool's.
They need the original question, verbatim, in the language it was asked. Not a summary, and not only a translation. The phrasing carries information: hesitancy, what the person already believes, which of several things they are actually worried about. A summary strips exactly that, and a translation alone removes the ability to check whether a misunderstanding started in the question rather than in the answer.
They need to know what the tool already said. A routed conversation where the employee has already received a partial answer is different from a clean refusal, and the receiving person needs to know which they are dealing with before replying. Contradicting something the employee has already been told, without realising they were told it, is the most avoidable bad outcome in this whole arrangement.
They need a queue, not a notification. Routed questions arriving as individual messages into somebody's inbox will be handled inconsistently, because they compete with everything else in that inbox and there is no view of what is outstanding. A small shared queue with the four refusal categories as tags is enough, and it makes the volume visible, which is what justifies the resourcing later.
They need permission to answer slowly. The routed categories are the hard questions by definition. If the receiving person is measured on the same response time as the procedural queue, they will rush exactly the questions that should not be rushed. Give the routed queue its own, longer, published target.
And they need a route onwards. Some routed questions will be beyond the receiving person too, particularly the jurisdiction-specific ones where the honest answer requires local advice. Name that second step when you name the first, or the refusal chain stops at somebody who cannot finish it either.
The practical test of whether you have set this up properly is to ask the receiving person what proportion of routed questions they could answer immediately. If it is low, the refusal rules are too broad or the handover is missing context, and both are fixable.
The Refusal in Every Language
Refusals are short, high-stakes content that gets read at the worst moment, which makes them the least suitable thing in your whole deployment for loose machine translation.
Three properties make them unusually sensitive. They are read by somebody who is already uncertain, so register matters more than usual and a refusal that lands as curt reads as a brush-off. They contain a name and a timeframe, which are exactly the elements that get mangled by automatic handling. And they are the one piece of content whose failure mode is the employee giving up rather than asking again.
So treat the four refusals as the top bucket: human translation, reviewed by somebody who knows both the language and what the refusal is for, with the names and the timings checked individually in each version. It is four short paragraphs per language. The whole job is an hour per language and it is the best-spent hour in the deployment.
Two specifics worth getting right. Make sure the stated response time is expressed in the reader's working days rather than head office hours in every version, which is easy to lose in translation because the original phrasing often assumes a shared context. And check that the name survives: a personal name rendered phonetically or a team name translated descriptively will not match anything the employee can search for, which turns a named route back into a generic one.
Then review them when anything about the route changes. A refusal naming somebody who left eight months ago is worse than no refusal, because it has consumed the employee's trust before failing.
Testing the Refusals Before Launch
Twenty minutes, before anybody sees the tool, and it catches most misconfigurations.
Write ten questions you want refused. Two per category, plus two deliberately borderline ones where you are genuinely unsure whether a refusal is right. The borderline pair is the most informative, because it tells you where the rules currently draw the line.
Ask each one in each supported language. Not just your working language. Refusal rules keyed to phrasing frequently fire in the language they were written in and fail in the others, and this is the single most common defect in a multilingual deployment.
Check four things in each response. That it refused rather than answered. That it gave a reason. That it named a real route. And that the routed conversation actually arrived, with the original question attached, wherever it was supposed to go. The fourth is the one people skip and the one that is most often broken.
Then ask ten questions you want answered. This is the control, and it matters because a refusal set drawn too broadly is a real cost. If any of the ten procedural questions gets routed, the rules are catching too much and the receiving person is about to be buried in things a document already covers.
Record the results and keep the question set. Rerun it after any content change, any configuration change and any vendor update. Behaviour in this category changes with model updates that nobody tells you about, and a twenty-question regression test run quarterly is the cheapest protection available against a tool that quietly starts answering things it used to decline.
How to Choose: Five Questions Before You Talk to Any Vendor
Ask it something it should refuse, in the demo. Not a hard question, a question requiring judgement about an individual. What it does is the most informative thing you will see in the whole call, and it is the test almost nobody runs because demos are steered towards capability.
Can refusals be keyed to topics rather than phrases? Phrase-based rules are brittle against the variety of real questions. Topic-level configuration is what makes the sensitive categories reliably catchable.
Does it hedge when the retrieval is a loose match? The behaviour you want is a qualification when the tool is working from adjacent rather than exact material. Uniform confidence regardless of match quality transfers all the risk to your review process.
Are routed conversations reported separately, with the reason? You cannot tune what you cannot see, and a routed question filed as an unresolved one will push you to configure fewer refusals rather than better ones.
Does the routed conversation carry its context? Ask to see the handover rather than hearing that one exists.
What the Tools Say About Their Own Limits
The most useful signal in this category is whether a vendor states its constraints without being asked. Each product below was checked on its own site on 5 and 7 October 2026, and each is described separately, because mixing vendors and figures in one paragraph is how a real price ends up attached to the wrong product.
Matram
Disclosure: Matram is owned by the same people who publish HROpsLab. It appears here because it competes in this category and is assessed against the same criteria as everything else on this page, with its limitations stated in the same detail.
Its stated limits are the three that this article says matter, and they are worth repeating plainly because this is an article about limits. There is no free tier, so it cannot be trialled indefinitely on a small population, although there is a 30-day trial that does not require a card. It answers questions and does not action tickets, so it explains a process rather than completing one. And it is not an HR workflow platform, so anything requiring a record to change is outside it.
What it does well for the subject of this article is that it is document-trained, which means a refusal can be written as content and retrieved like any other answer, rather than depending on a rule somebody configured in a console. It publishes 95+ languages and prices flat at $29, $69 and $199 a month with seats unlimited, which matters here because refusals route to people and routing to more people should not cost more.
SiteGPT
Document-trained, so refusals can be written as content in the same way. It publishes the same 95+ languages figure, so neither it nor the above leads on reach. Priced at $468 and $948 billed yearly, with the monthly-equivalent figures applying only on that annual commitment. Starter covers one chatbot and a capped page count.
Chatling
Document-trained with a free tier, which makes it the cheapest way to test whether your written refusals are actually retrieved in practice before committing. Its published language figure is inconsistent, 80+ in one homepage section and over 85 in another.
Botsonic
Publishes 50+ languages, plainly stated, which is narrower than the leaders and a figure you can rely on.
Crisp
A broader support platform with live chat as well as an answer layer, which has a genuine advantage for this topic: a refusal can route into a human conversation in the same interface rather than into a different system. Priced per workspace at Free, $45, $95 and $295 a month, so adding the people who receive routed questions does not change the bill. No published language count.
Zendesk
Suite Team is $55 per agent per month with the Copilot add-on a further $50 per agent per month. It has the deepest routing and workflow capability here, which is the relevant strength, and the per-seat basis means every person who receives routed questions is a licence. No published language figure, and the pricing page geo-redirects.
Freshservice
Tiered per agent at $19, $49 and $99 per month, with Freddy AI priced separately at $29 per agent per month. Covers workflow as well as answers, so it can action some of the third category rather than only explaining it. No published language count.
Three vendors publish no price at all. Leena AI and Moveworks are demo-only, and Moveworks has no working pricing page. Tidio has replaced plan pricing with a usage calculator, so its on-page figures are conversation counts rather than money.
The Comparison
| Tool | Refusals writable as content | Can action requests | Language count published | Pricing basis |
|---|---|---|---|---|
| Matram | Yes | No, answers only | 95+ | Flat, seats unlimited |
| SiteGPT | Yes | No, answers only | 95+ | Annual per plan |
| Chatling | Yes | No, answers only | Inconsistent, 80+ and over 85 | Per plan, free tier |
| Botsonic | Yes | No, answers only | 50+ | Per plan |
| Crisp | Partially | Routes into live chat | None published | Per workspace |
| Zendesk | Partially | Yes, with workflow | None published | Per agent plus AI add-on |
| Freshservice | Partially | Yes, with workflow | None published | Per agent plus AI add-on |
The Decision Table
| Situation | Scale | Setup | Primary Pain | Recommended Starting Point |
|---|---|---|---|---|
| All questions procedural, nothing individual | Any | Default configuration | None. Answering everything is fine | Deploy and move on. Revisit if the mix changes |
| Mostly procedural with a sensitive minority | Under 200 | Four written refusals | Tool answers judgement questions confidently | Write the four refusals before launch, not after |
| Multi-country, local documents incomplete | 200 to 1,000 | Refusal per divergent topic | General answers delivered unqualified | Write the jurisdiction refusals first. They are the highest value |
| Requirement is to process requests, not explain them | Any | Workflow platform | An answer bot explains but cannot complete | Shortlist workflow tools. An answer bot is the wrong purchase |
| Disclosures arriving through the support channel | Any | Conservative routing rule | Sensitive topics handled by a tool | Route the whole category immediately, with no answer attempted |
| Restructure or change programme underway | Any | Refusals plus named route | Anxious individual questions arrive in volume | Configure before announcing. The volume arrives on day one |
| Reporting counts routed questions as failures | Any | Metric change | Configuration pushed towards answering everything | Report routed conversations as a separate successful outcome |
The Two Things That Look Like Refusals and Are Not
Worth separating out, because both get counted as good behaviour and neither is.
The non-answer that gives no route. A reply saying the tool cannot help with that, full stop. This is technically a refusal and it is functionally a dead end, and it is the default fallback text in most products. It is worse than a wrong answer in one specific way: a wrong answer at least gives the employee something to check with a colleague, whereas a routeless refusal gives them nothing and teaches them the channel is useless. Replace every generic fallback with a routed one before launch, including the ones you do not expect to fire.
The hedge that answers anyway. A reply that notes the answer may vary by circumstances and then provides the general answer in the next sentence. This is the most insidious pattern in the category because it looks responsible and behaves exactly like an unqualified answer. People read the specific content and discount the caveat, which is how everybody reads everything. If the question needed a refusal, the reply should not contain the general answer at all.
Both of these pass a casual transcript review, because the first contains an appropriate admission and the second contains an appropriate caveat. Finding them requires asking a different question of each transcript: did this employee end up with an answer they could act on, and should they have.
A third, rarer, is worth a mention. Some tools respond to a sensitive question by asking a clarifying question, which delays the refusal by one turn and extracts more personal detail along the way. That is a bad outcome even if the eventual routing is correct, because the employee has now disclosed more into a channel that was never the right place for it. Check for it specifically on the disclosure category.
Measuring Whether the Refusals Are Right
Two failure directions, and you need to watch both.
Under-routing shows up as answers nobody would have given. Find them by sampling transcripts on your sensitive topic list specifically rather than at random, because the whole point is that these are a small proportion of volume. One pass of twenty conversations on those topics will tell you whether the rules are firing.
Over-routing shows up as a rising proportion of routed questions that the receiving person answers with something the tool already had in a document. That is a refusal rule drawn too broadly, and it costs you the thing you bought the tool for. Ask whoever receives routed questions to flag these, because they are the only people who can see them.
The useful ratio is routed questions as a proportion of total, tracked over time and broken down by reason. A healthy deployment usually settles somewhere in single figures. A number climbing month on month without a change in the question mix means the rules are broadening, and a number near zero on a question set that includes individual circumstances means they are not firing at all.
Also watch what happens after a refusal. A routed conversation that the employee abandons rather than pursuing is a handover that failed, and the usual cause is a generic destination or no stated timing. That is fixable in the refusal wording rather than in the routing.
Break all of these down by language, which is the part most reporting makes awkward and which matters most. Refusal rules behave differently across languages, because the mechanism that recognises a sensitive topic is doing so on text, and the text is different. A deployment where the refusal rate in your working language is six per cent and in your second language is one per cent does not have a cultural difference in what people ask. It has rules that are not firing, and the population with the weakest ability to challenge an answer is the one receiving the most unqualified ones.
One further measure is worth adding after a few months: the proportion of routed questions that the receiving person resolves without escalating further. A high figure means the refusals are catching the right things and the handover is working. A low one means you are routing questions to somebody who cannot answer them either, which is a refusal chain with no end and is experienced by the employee as being passed around.
What Getting This Wrong Costs
The direct cost is a decision made on a synthesis nobody authorised. Usually recoverable, occasionally not, and the awkward feature is that you will often never know it happened, because the employee had no reason to mention a question they considered answered.
The second cost is positional. Once a tool has answered a judgement question in a particular way, repeatedly, the company has effectively taken a position, and reversing it with an individual is harder than never having stated it. People reasonably treat a written answer from a company system as the company's answer, and they are not wrong to.
The third cost is trust running in both directions. Employees who receive confident answers on sensitive topics and later discover the answers were not authoritative stop using the channel for anything that matters, which leaves it handling only the questions that were never difficult. And the HR team loses the thing a well-configured tool actually provides, which is a reliable filter that brings them the questions that need them, with context attached.
There is also a cost to the people receiving routed questions, which is worth planning for rather than discovering. A well-configured tool sends them a queue consisting entirely of the difficult questions, with the easy ones removed. That is the correct design and it changes the nature of the work substantially, from a mixed queue with some straightforward wins to a concentrated stream of judgement calls. Resource it accordingly and say so openly, because the alternative is a quiet complaint that the new tool has made somebody's job harder, which is both true and not a reason to configure it worse.
So the question to take into a demo is not how much the tool can answer. It is what happens when you ask it something it should not.
When You're Ready to Move Beyond Answering Everything
Most deployments start configured to answer as much as possible, and that is not a mistake so much as a default. The tool arrives optimised to demonstrate capability, the implementation focuses on coverage, and nobody writes refusals because refusals are not what anybody bought. The first few months look good on the metrics that were set up.
What makes it stop working is the first sensitive question, which arrives much earlier than people expect and is usually not recognised at the time. The shift required is small: it is accepting that the measure of a good deployment is not the proportion of questions answered but the proportion answered correctly plus the proportion routed correctly, which are different numbers and only the first is on most dashboards.
The sequence is short enough to do in a week. List the topics where the answer depends on individual circumstances, which you probably already have from a handbook audit. Write four refusals, one per category, with a named route and a stated response time, and translate them properly into every language you support, because these are the sentences you least want machine-translated loosely. Configure the sensitive-disclosure rule conservatively. Change the reporting so a routed question counts as a success. Then ask the tool something it should refuse, and check.
Frequently Asked Questions
What kinds of HR questions should a chatbot not answer?
Four categories need a person regardless of how capable the tool is. Anything whose correct answer depends on the individual's contract, tenure, location or situation. Anything jurisdiction-specific where you have no local document, because the tool will answer from the general one unqualified. Anything requiring an action rather than an explanation, such as changing an election or correcting a record. And anything where somebody is disclosing a circumstance rather than asking a question, which includes health, conduct and safety matters that frequently arrive phrased as questions.
Is a chatbot refusing to answer a failure?
No, and treating it as one is why many deployments are configured badly. A refusal that explains why and names the person who can help is a successful outcome and should be measured as one, because it routes a question requiring judgement to somebody who can take responsibility for the answer. If your reporting counts routed conversations as deflection failures, the metrics will push you to configure fewer refusals rather than better ones, which is the opposite of what you want.
Why does answering rather than actioning matter?
Because it is the limit most often misunderstood at the point of purchase. An answer bot can explain exactly how to change a benefits election and cannot change it, and a demo showing the first looks very similar to a demo of the second. If your actual requirement is that requests get processed rather than explained, you are shopping for a workflow platform and an answer bot will disappoint regardless of how good its answers are.
Can we fix this by training the tool on more documents?
It helps with coverage gaps, which are real and worth fixing, and it does not help with the four categories. Those questions are not answerable from documents at all, because they require somebody to exercise judgement or to take an action, so adding material makes the tool more confident in precisely the areas where confidence is the problem. The fix is written refusals and topic-level routing rather than more content.
How should a refusal be worded?
Four parts: say you cannot answer and why in one sentence, name a specific person or monitored inbox rather than "contact HR", state when they will hear back in their own working days, and carry the original question across so the employee does not have to retype their situation. The explanation matters because it distinguishes a deliberate refusal from a failure to understand, and the named route matters because a handover to a generic destination is where routed questions go to wait.
How do we know if the refusals are configured correctly?
Watch both failure directions. Under-routing shows up as answers nobody at the company would have given, found by sampling transcripts specifically on your sensitive topics rather than at random. Over-routing shows up as routed questions that the receiving person answers using something the tool already had in a document, which only the receiving person can see, so ask them to flag it. Track routed conversations as a proportion of total with reasons attached, and treat a number near zero on a question set involving individual circumstances as evidence the rules are not firing.
Is it better to pick a vendor that admits its limits?
Yes, and it is one of the more reliable signals available when comparing products that otherwise look similar. The limits exist in every product in this category, so a vendor stating them plainly has demonstrably thought about the handover, which is the part of the deployment that determines whether a refusal is useful or merely a dead end. A vendor whose materials imply the tool handles everything has either not considered the sensitive categories or has decided not to mention them, and neither is reassuring.
HROpsLab takes no vendor money and publishes no paid placements, which is why the owner's own product is listed above with its three stated limits rather than its three best features.