AI HR Tools 25 min read

Multilingual HR Chatbots: What a 95-Language Claim Actually Covers

Only four of eight vendors publish a language count, two publish the same figure, and one contradicts itself on its own homepage. What to check before you buy.

Rachel Kim Rachel Kim • • 25 min read

TL;DR

  • A multilingual HR chatbot answers employee policy questions in more than one language, usually by reading your own documents and replying in whatever language the question arrived in.
  • If everybody on your payroll works comfortably in one language, you do not need this. You need a good single-language answer bot, or possibly just a better handbook.
  • Only four of the eight tools checked for this piece publish a language count at all. Two of those four publish the same number, and one contradicts itself in two places on its own homepage.
  • A language count measures the model's translation reach. It does not measure whether your policy documents exist in that language, which is the part that decides whether the answer is right.
  • Coverage and pricing collide. Adding a market usually means adding people, and most support tools still charge per agent or per seat, so the language decision and the budget decision are the same decision.
  • Test a claim before you believe it. Ask the same awkward policy question in three languages and read all three answers side by side, which takes an afternoon and settles more than any demo.

The Demo That Answered in Polish and Then Quietly Guessed

A 340-person company with staff in six countries sat through four vendor demos in a fortnight. Every one of them handled the language question the same way. Somebody typed a question in Polish, the tool answered in Polish, and the room nodded. The HR director, who does not read Polish, nodded too. So did the two people from IT.

A week later somebody on the HR team asked a Polish-speaking colleague to read the transcript properly. The grammar was good. The answer was wrong. The question had been about notice during a probation period, the handbook had a general section and no Polish-specific one, and the tool had answered from the general section without ever saying that it was doing so. It read as authoritative because it was fluent. Nobody in the demo could tell the difference, because fluency is the only thing a non-speaker can assess.

That is the whole problem in one transcript. The language count on the pricing page was accurate. The tool really did handle Polish, and it would have handled another ninety-odd languages just as fluently. What no vendor advertises is that a language count describes the model's ability to translate, not your organisation's ability to supply a correct answer in that language, and those two things fail independently. A tool with perfect Polish and no Polish source material will produce a confident wrong answer every time, and it will sound exactly like a right one.

This is what a multilingual HR chatbot is supposed to solve, and it is worth being precise about which half of it actually gets solved by buying something.

When You Don't Actually Need a Multilingual HR Chatbot

When the manual way is genuinely fine

Everybody on the payroll works in one language, the handbook exists in that language, and HR answers about fifteen questions a week. At that size the honest advice is to improve the handbook and leave the tooling alone. A multilingual tool adds a vendor relationship, a monthly bill and a new review burden, and it removes nothing you are currently struggling with. The threshold is not headcount. It is whether anybody is answering questions in a language your handbook does not cover.

When friction starts appearing

The first signal is rarely a complaint. It is a pattern where the same question arrives repeatedly from the same country, phrased slightly differently each time, because the person asking cannot find the answer and is not confident enough in their second language to be sure they have understood the handbook. That is a translation gap presenting as a volume problem. It is cheaper to fix with two translated documents than with a platform, and fixing it that way tells you whether the volume was ever really about language.

Free Weekly Briefing Stay ahead of what's changing in HR and people ops.

Join 4,200+ leaders getting practical insights every week — no fluff, just signal.

Join Free →

When it becomes a liability

The cost of not having this exceeds the cost of having it at the point where somebody acts on an answer that was wrong in their jurisdiction. Not a legal catastrophe necessarily, just an employee who believed something about notice, or leave, or a benefit, because a system told them, and made a decision on that basis. The exposure is not the chatbot. It is that nobody can reconstruct what the employee was told, because the answer was generated in a language nobody on the HR team reads.

The edge case that forces it

A rapid expansion into two or three markets at once, or an acquisition that brings a population who work in a different language. Both produce a step change rather than a gradual one. A process that absorbed one Spanish-speaking joiner a quarter does not survive forty of them arriving in March, and the people best placed to notice the answers are wrong are precisely the new joiners who have least standing to say so.

Five Questions People Ask Before They Shortlist Anything

"How many languages do we actually need?" Almost always fewer than the vendor list implies, and the right way to get the number is to count people rather than countries. Pull your headcount by working language, not by office location. A company with staff in nine countries routinely finds that seven of those populations work in English day to day and two do not, which makes it a two-language problem with a long tail, not a nine-language one. This matters because every extra language you commit to supporting is a document set somebody has to maintain.

"Will it answer in the language the question came in?" Usually yes, and this is the part that works well. Language detection on an inbound message is close to a solved problem, and the major tools all do it without configuration. The question worth asking instead is what happens when detection is ambiguous, which it is for short messages. A three-word question in a language that shares vocabulary with another can be misdetected, and the useful vendor behaviour is to ask rather than guess.

"Does it need our documents translated first?" This is the question that separates the tools, and most buyers do not ask it. Some systems read a source document in one language and answer in another, translating at the point of reply. Others expect a document set per language. The first is far less work and introduces a translation step you cannot see or review. The second is more work and produces an answer you can actually audit. Neither is wrong, but they imply completely different amounts of ongoing effort.

"What happens when the answer differs by country?" The honest answer from most vendors is that it does not handle this unless you structure your documents so that it can. Language and jurisdiction are not the same axis, and conflating them is the most common design error in this whole category. Spanish is spoken in a lot of places with very different rules. A tool that routes on language alone will cheerfully give a Mexican employee an answer written for Spain.

"Can we see what it told somebody?" Ask for the transcript export in the first call, and ask whether the stored record includes the original question, the answer, and the source document the answer came from. Without the third of those you cannot review anything, because you cannot tell whether a wrong answer was a translation failure or a document failure. The distinction decides whether you fix the tool or fix the handbook.

What a Language Count Is Actually Measuring

A number on a pricing page is doing one of three jobs, and vendors rarely say which.

The first is model reach. The underlying language model handles a given language at some level of competence, so the vendor lists it. This is the cheapest kind of claim to make and the most common. It is not dishonest, but it tells you about the model rather than about the product, and the model is the same one several of these vendors are using.

The second is interface coverage. The widget, the buttons, the fallback messages and the error states have been translated. This is a real engineering commitment and a much smaller number, usually a dozen or two. It is also the thing an employee notices first, because a reply in fluent Polish wrapped in an English interface with an English "Was this helpful?" prompt reads as half-finished.

The third is tested coverage, meaning somebody at the vendor has actually evaluated answer quality in that language against a known-correct set. Almost nobody publishes this, and when you ask, the answer is usually about the model again. The gap between the first claim and the third is where the demo in Lisbon went wrong.

So when you read a count, work out which of the three you are being told. Then ask for the other two.

The Three Ways Vendors Build Language Coverage

Translate at the point of reply

The tool keeps one source document set, usually in English, and translates the answer as it generates it. This is how most of the general-purpose answer bots work. It is right when your documents are genuinely universal, your populations are small, and the questions are about things that do not vary by country, like how to book leave in a system everybody uses. It fails when the source document has no section covering the asker's situation, because the translation layer has nothing to tell you that. The answer arrives fluent and unqualified.

Maintain a document set per language

You supply the handbook in each language, and the tool answers from the matching set. This is more work up front and considerably more work forever, because every policy change becomes several document changes. It is right when the answers genuinely differ by population and when you need to be able to show what a given employee was told and where it came from. It fails on staleness. A document set per language is a maintenance commitment, and the second-language versions are always the ones that fall behind.

Route to a person

Not a tool strategy so much as the absence of one, but it belongs on the list because it is often the correct answer for the hard questions.

Approach Right when Fails when Ongoing effort
Translate at the point of reply Documents are genuinely universal and questions do not vary by country The source has no section for the asker, so the answer arrives fluent and unqualified Low, and the translation step is invisible to you
Document set per language Answers genuinely differ by population and you need to show what somebody was told Staleness. Second-language sets always fall behind High, and it compounds with every policy change
Route to a person Small populations in complicated markets Cost and coverage hours, and it becomes a hiring requirement Low content effort, real staffing effort The bot handles the high-volume universal questions in any language, and anything touching a jurisdiction-specific topic goes to somebody who can handle it properly. This is right for small populations in complicated markets. It fails on cost and on coverage hours, which is a different article, and it means your language requirement is now a hiring requirement.

How to Choose: Five Questions Before You Talk to Any Vendor

Which languages do you actually have, by headcount rather than by flag? Write the list down before any demo. Two columns, language and number of people. The list is almost always shorter than expected and it changes the conversation completely, because you stop comparing vendors on a total and start asking whether they are good at your four.

Does your handbook already exist in those languages? If it does not, you are buying a translation layer whether you meant to or not, and you should decide that deliberately. If it does, ask every vendor how it handles multiple source sets, because some handle it badly.

Do your answers vary by jurisdiction, or only by language? Be honest here. Most companies have a handful of topics that genuinely diverge and a long list that does not. Identify the divergent ones, because they are the only ones where this decision carries any real risk, and they are what you should test in the demo.

Who will read the answers? Somebody has to review output in a language they speak, or the quality question is permanently unanswerable. If nobody internally speaks your second language, that is a finding, and it should push you towards a tool whose answers are auditable rather than one whose answers are merely fluent.

What is the pricing basis? Per seat, per workspace, per resolution or flat. This sounds like a procurement detail and it is actually the decision, because the basis determines whether expanding into a fifth market costs you nothing or costs you a line item per person. Work out the twelve-month figure at your headcount in two years, not today.

Eight Options and What Each One Publishes

A note on language counts before the list, because it is the most striking thing about this category. Only four of these eight publish a language figure at all. Each was checked against the vendor's own site on 7 October 2026. Two of the four publish the same number as each other. One of them publishes two different numbers in two sections of the same homepage. The other four describe multilingual support without attaching a count to it, which means there is no figure to compare and no way to hold them to one.

Matram

Disclosure: Matram is owned by the same people who publish HROpsLab. It appears here because it competes in this category and is assessed against the same criteria as everything else on this page, with its limitations stated in the same detail.

Best for: teams whose language requirement is growing faster than their support headcount, and who want the bill to stop tracking the number of people who answer questions.

Why companies choose it: it publishes 95+ languages, and it prices flat per month rather than per seat, at $29, $69 and $199, with seats unlimited on every tier. That combination is the point. If your reason for needing more languages is that you have hired in more places, a per-seat tool charges you twice for the same expansion, once in salary and once in licence. It trains on your own documents, so the answers come from your handbook rather than from general knowledge, and there is a 30-day trial with no card required, which makes the afternoon test described below cheap to run.

Where it struggles: there is no free tier, so you cannot park it on a small population indefinitely to see what happens. It answers questions and does not action tickets, so if your requirement is "and then it should process the leave request", that is a different purchase. It is not an HR workflow platform and does not pretend to be one. And its 95+ figure is matched exactly by SiteGPT below, so on language reach alone there is nothing to choose between them.

SiteGPT

Best for: teams that want a document-trained answer bot and are comfortable with annual billing.

Why companies choose it: it publishes the same 95+ languages figure, which makes it the honest co-leader on reach rather than a runner-up. It is built around training on your own site and documents, so it sits in the same part of the market as Matram and solves the same core problem.

Where it struggles: the pricing is annual-first. Starter is $468 and Growth is $948, both billed yearly, and the monthly-equivalent figures that appear alongside them are only available on that annual commitment. Starter covers one chatbot and a capped page count, which is tight if you were planning a separate bot per population. Quoting the monthly equivalent without the annual qualifier is the single easiest way to misbudget this one.

Chatling

Best for: teams that want to try a document-trained bot without a commitment, because there is a free plan.

Why companies choose it: low barrier to entry, and a straightforward setup aimed at people who are not going to configure an intent tree.

Where it struggles: it cannot keep its own language figure straight. One section of its homepage says 80+ and another section of the same page says over 85. That is not a large discrepancy and it is not evidence of anything sinister, but it is the clearest possible illustration of how lightly these numbers are held. If a vendor has not reconciled the figure across one page, treat the figure as marketing rather than specification.

Botsonic

Best for: teams already using the wider product suite it belongs to.

Why companies choose it: it publishes a concrete figure, 50+ languages, which is lower than the leaders and stated plainly. A smaller number you can trust is more useful than a larger one you cannot.

Where it struggles: the coverage is genuinely narrower, so check your specific list against it rather than assuming the long tail is handled. Its product URL has also moved, which is worth knowing if you have old bookmarks or an old vendor record.

Crisp

Best for: teams that want a shared inbox and live chat with an answer bot attached, rather than an answer bot on its own.

Why companies choose it: the pricing basis is unusual and genuinely advantageous at scale. It charges per workspace rather than per seat, at Free, $45, $95 and $295 a month, which means adding people to the team does not change the bill. For a growing support function that is a meaningful structural difference.

Where it struggles: it publishes no language count at all. Multilingual capability is described rather than quantified, so there is no figure to hold it to and no way to compare it on this axis. It is also a broader product than an answer bot, which is a benefit if you want the inbox and overhead if you do not.

Chatbot.com

Best for: teams that want a visual builder and are happy to design flows explicitly.

Why companies choose it: predictable, conventional licensing at $19 per user per month, and a mature builder for people who want to control exactly what the bot says rather than letting it generate from documents.

Where it struggles: per-user pricing is precisely the model that punishes geographic expansion, because every market you enter tends to come with the people who support it. It publishes no language figure either, so the multilingual question has to be answered in a call.

Zendesk

Best for: organisations already running Zendesk, where the answer bot is an add-on decision rather than a platform decision.

Why companies choose it: it is the incumbent in a lot of support functions, and extending something you already run is a shorter project than introducing something you do not.

Where it struggles: the cost structure is the hardest in this list for a multi-country team. Suite Team is $55 per agent per month and the Copilot add-on is a further $50 per agent per month, so the per-person figure roughly doubles once you want the AI layer. It publishes no language count. And its pricing page geo-redirects, so confirm which currency you are looking at before building a model on it.

Freshservice

Best for: teams that want service management and employee support in one place.

Why companies choose it: tiered per-agent pricing at $19, $49 and $99 a month gives a low entry point, and the platform covers workflow as well as answers, which the pure answer bots do not.

Where it struggles: the AI layer is a separate line, with Freddy AI at $29 per agent per month on top of the tier. It publishes no language figure. And like every per-agent product here, the model works against you specifically in the scenario that drove you to look at multilingual tooling in the first place.

Two others are worth naming and cannot be compared on price, because they publish none. Leena AI and Moveworks are both demo-only in this space. Moveworks does not even have a working pricing page. Tidio has abandoned fixed plan pricing entirely in favour of a usage calculator, so the numbers on its page are conversation counts rather than money, which is a trap worth knowing about before you screenshot it for a comparison deck.

What Each One Published

Tool Language count published Pricing basis Published price
Matram 95+ Flat per month, seats unlimited $29, $69, $199
SiteGPT 95+ Annual, per plan $468, $948 billed yearly
Chatling 80+ in one place, over 85 in another Per plan, free tier Free tier available
Botsonic 50+ Per plan Not compared here
Crisp None published Per workspace Free, $45, $95, $295
Chatbot.com None published Per user $19
Zendesk None published Per agent plus AI add-on $55 plus $50
Freshservice None published Per agent plus AI add-on $19, $49, $99 plus $29

Read the first column as the finding. Four blanks and one self-contradiction out of eight, in a category where multilingual capability is the headline feature, tells you how little weight the industry puts on a figure it expects nobody to check.

The Decision Table

Situation Scale Setup Primary Pain Recommended Starting Point
One working language, handbook already good Any None needed Questions are repetitive, not multilingual Fix the handbook. Revisit tooling in a year
Two or three languages, small second-language populations Under 200 Document-trained bot, one source set Second-language staff cannot find answers A document-trained bot with a trial, tested in all three languages before committing
Growing headcount across many markets, support team growing with it 200 to 1,000 Flat-priced answer bot Per-seat licensing compounds every expansion Matram, on the seat economics, with the language parity against SiteGPT understood
Document-trained bot wanted, annual budget cycle 200 to 1,000 Annual commitment Needs one tool, predictable yearly cost SiteGPT, budgeted on the annual figure rather than the monthly equivalent
Live chat and shared inbox needed as well as answers Any Workspace-priced platform Team size drives cost on per-seat tools Crisp, on the per-workspace basis
Already running a service management platform 500 plus AI layer on the incumbent Integration effort outweighs licence saving Extend the incumbent, with the AI add-on costed per agent
Answers genuinely differ by country on key topics Any Document set per jurisdiction Language routing gives the wrong country's answer Restructure documents first, then shortlist. No tool fixes this for you

How to Structure Documents So the Tool Can Get Jurisdiction Right

Two sections above told you to fix the documents before buying anything, which is easy to say and vague enough to ignore. Here is what it actually involves, because it is a smaller job than it sounds and it is the only part of this that no vendor can do for you.

Start by splitting your handbook into two piles rather than translating it. The first pile is everything that is true for everybody: how to book leave in the system, who to contact about a payslip, what the expenses process is, how the review cycle runs. This is usually most of it, and it is the pile that benefits from translate-at-reply, because there is one correct answer and the only question is which language it arrives in.

The second pile is the topics where the answer changes depending on where somebody is employed. Notice, probation, statutory leave entitlements, sick pay, public holidays, anything touching termination. This pile is nearly always short, often eight or ten topics, and it is where every expensive mistake in this category happens. Write these as separate documents per population rather than as one document with a table of exceptions, because a tool retrieving from a long document with a regional table buried in it will frequently return the general paragraph and stop.

Then do the step most teams skip: put the scope in the text of the document itself, in the first line, in plain words. Not in the filename, not in a folder name, not in a metadata field. "This section applies to employees engaged in Germany" as a sentence at the top of the document means the retrieval layer has the scope in the same place as the content, so an answer drawn from it carries the qualification into the reply. Filenames and folders are invisible to most retrieval, which is why a well-organised document library still produces unscoped answers.

Finally, write a short refusal into the second pile. One paragraph in each jurisdiction-sensitive document saying that the specifics depend on individual circumstances and pointing at a named person or inbox. This sounds like a hedge and it is doing real work: it gives the tool something correct to retrieve for the questions it should not answer confidently, which is far better than leaving it to improvise from the general pile.

Done properly this is a two-day job for one person who knows the handbook, and it makes every subsequent tool decision easier. It also tells you something useful on its own. If the second pile turns out to be long rather than short, your problem is policy fragmentation rather than language, and a chatbot will not touch it.

Testing a Language Claim in an Afternoon

This is the part that costs nothing and almost nobody does.

Pick three questions. One should be universal, like how to request leave in whatever system you use. One should be jurisdiction-sensitive, something where you know the answer differs between two of your markets. One should be deliberately ambiguous, a short badly-phrased question of the kind people actually send.

Ask all three in each of your real languages, in each tool's trial. Then get a speaker of each language to read the answers, and ask them two things rather than one: is this fluent, and is this correct. Those come apart more often than you would expect, and the gap between them is the only number in this whole exercise that describes your situation rather than the vendor's.

Pay particular attention to the jurisdiction-sensitive question. The failure mode you are hunting is not a bad translation. It is a confident answer drawn from a document that does not apply to the person asking, delivered without any signal that a general source was used. A good tool will hedge or escalate. A weak one will answer.

Finally, ask the ambiguous question and watch what the tool does with detection. Asking for clarification is the right behaviour. Guessing and proceeding is the behaviour that produced the Polish transcript.

What Getting This Wrong Costs

The direct cost is the licence, and it is the smallest of the three. A per-agent tool at an AI-inclusive rate runs to a four-figure monthly sum for a mid-sized support team, and a flat-priced one does not, and that difference is arithmetic you can do in a spreadsheet in ten minutes.

The second cost is the maintenance you did not plan for. A document set per language is an ongoing obligation, and the obligation is invisible at purchase because the first translation happens during implementation when everybody is paying attention. It is the third policy change eighteen months later, applied to the English handbook and not to the other four, that creates the real exposure. By then nobody remembers that the other four exist as separate files.

The third cost is trust, and it is the one that does not show up in any budget. An employee who receives a wrong answer in their own language learns something specific: that the system is confident and unreliable. They do not escalate, because escalating in a second language to a team that works in another one is effortful. They just stop asking, and route around HR entirely by asking a colleague. Your ticket volume falls, which looks like success.

So the question to carry into the vendor calls is not which tool handles the most languages. It is whether you will be able to tell, six months after launch, that an answer was wrong.

When You're Ready to Move Beyond English Plus a Translation Tab

Most companies arrive here from the same place. The handbook is in English, somebody on the HR team pastes awkward questions into a translation tool, and it works well enough that nobody has escalated it. That arrangement is more defensible than it sounds, and it is worth saying plainly that it is not an emergency.

What makes it stop working is volume and consistency rather than quality. Two people doing ad hoc translation produce two different answers to the same question, and neither is recorded anywhere, so the inconsistency is undetectable until somebody compares notes. The case for a tool is not that it translates better than the translation tab. It is that it answers from one source and keeps a record, which is the thing a human workaround structurally cannot do.

The honest sequencing is documents first, tool second. Work out which of your answers genuinely differ by population, get those into shape, and only then shortlist. A tool bought before that work is done will faithfully reproduce whatever ambiguity your handbook contains, in ninety-five languages, very fluently.


Frequently Asked Questions

What is a multilingual HR chatbot?

It is a tool that answers employee questions about policy, benefits, leave and process in more than one language, usually by reading your own handbook and internal documents and then replying in whichever language the question was asked in. The distinguishing feature against a general chatbot is that the answers are supposed to come from your material rather than from the model's general knowledge, which is what makes the answer specific to your company rather than to companies in general.

How many languages do we actually need?

Fewer than any vendor list suggests, and the way to find out is to count your own people by the language they work in rather than by the country they sit in. Companies with staff in eight or nine countries commonly discover that most of those populations work in English day to day and only two or three do not, which turns a long list into a short one. Every additional language is a document set somebody has to keep current, so the number should be as small as the truth allows.

Is a higher language count better?

Not on its own, because the count usually describes what the underlying language model can translate rather than what the vendor has tested or what your documents cover. A tool publishing 50+ languages that handles your four well is a better purchase than one publishing a far larger number that has never been evaluated in any of them. Ask which of the three things the number measures, model reach, interface translation or tested answer quality, and you will usually find it is the first.

Do we need our handbook translated before we start?

It depends on which approach the tool takes, and this is the question most worth asking in a first call. Some tools keep a single source document set and translate the answer as they generate it, which is far less work and gives you an answer you cannot easily audit. Others expect a document set per language, which is more work up front and permanently, and produces an answer you can trace to a source, so the right choice depends on whether your answers genuinely differ between populations.

What is the difference between language and jurisdiction?

Language is how somebody asks, and jurisdiction is which rules apply to them, and treating the two as the same axis is the most common design error in this category. Spanish is spoken across many countries whose employment arrangements differ substantially, so a tool that routes purely on detected language will happily serve one country's answer to another country's employee. Where your answers genuinely diverge, the fix is in how you structure your documents rather than in which tool you buy, and the details of what applies to whom differ by jurisdiction and by circumstance, so local advice is the right route for anything consequential.

Which vendors publish a language count?

Of the eight checked on 7 October 2026, four publish a figure and four do not. Matram and SiteGPT both publish 95+, which makes them co-leaders rather than one leading, Botsonic publishes 50+, and Chatling publishes two different figures in two sections of the same homepage. Crisp, Chatbot.com, Zendesk and Freshservice describe multilingual support without attaching a number to it, so on this axis there is nothing to compare and the question has to be asked directly.

Why does pricing basis matter so much for multilingual support?

Because the thing that drove you to need more languages is usually the thing that adds people, and most support tools still charge by the person. Entering a new market tends to mean hiring somebody there, and on a per-agent or per-user product that shows up as a licence line on top of the salary, so the same expansion is billed twice. A flat or per-workspace basis breaks that link, which is why the pricing model deserves as much attention in this category as the feature list.

HROpsLab takes no vendor money and publishes no paid placements, which is why this page says "none published" four times rather than estimating.

Share on X Share on LinkedIn

What to do next?

Explore More Articles

Dig deeper into HR Ops strategy, tools, and workflows built for real teams.

Browse the blog →
Join the HROpsLab Community

Connect with People Ops practitioners sharing real workflows, tools, and challenges.

Join now →