HR Operations 25 min read

Who Signs Off on an HR Answer Nobody in HR Can Read

The review problem nobody plans for: approving a policy reply in a language your approver does not speak, and four ways teams actually solve it.

James Carter James Carter • • 25 min read

TL;DR

  • If your support tool answers in languages your HR team does not read, nobody is checking those answers, and fluency is the only signal available to the people who approve them.
  • If everybody on the HR team reads every language you support, you do not need this. You need a normal quality process and a sample size.
  • Fluent and correct come apart more often than people expect. A wrong answer in good grammar reads exactly like a right one, which is why this failure is silent rather than loud.
  • Four workable arrangements exist: a named internal reviewer per language, a reciprocal arrangement with the business, a paid external reviewer, and restricting what the tool is allowed to answer at all.
  • Review a sample, not everything. Twenty answers per language per quarter will find a systematic problem, and systematic problems are the ones that matter.
  • Make source attribution a shortlisting requirement. Without knowing which document an answer came from, you cannot tell a translation failure from a policy gap, and those have different fixes.

The Answer Everybody Approved By Nodding

A people operations team launched an answer bot across five languages. The launch review was thorough by most standards. They tested forty questions, read the answers, graded them, and signed off. The transcript file went into a shared drive and somebody set a reminder to review it quarterly.

Of the forty questions, thirty-one had been asked and answered in English. The other nine were in the four other languages, and in each case the reviewer had done the only thing available to them: pasted the answer into a translation tool, read the English version, and judged it reasonable. Which it was. The English round-trip of a wrong answer is a reasonable-sounding English paragraph, because the translation layer faithfully renders whatever it is given, including the error.

Four months later a Spanish-speaking employee mentioned in passing that the bot had told her something about probation that her manager said was not right. It had not invented anything. It had answered from the general handbook section because there was no Spain-specific one, and it had done so without qualification, in every one of those four months, to everybody who asked. The review process had been real, documented and diligent, and it was structurally incapable of catching the one class of error it most needed to catch, because the reviewer was assessing fluency and calling it accuracy.

This is the part of multilingual support that nobody puts on a slide, and it is the part that decides whether the whole arrangement is trustworthy.

When You Don't Actually Need a Review Process

When the manual way is genuinely fine

Everybody who approves HR answers reads every language the tool answers in. This is more common than it sounds at small scale, particularly in companies operating across two closely-related language markets where the HR team is genuinely bilingual. If that is you, you still need a quality process, but it is an ordinary one and this article is about a problem you do not have.

When friction starts appearing

The first signal is a reviewer asking for help. Somebody on the team says they are not sure whether an answer is right, or starts routing transcripts to a colleague in another department to check. That is the process telling you it has hit its limit, and it is the best possible moment to formalise it, because the person raising it has already identified the gap for you.

Free Weekly Briefing Stay ahead of what's changing in HR and people ops.

Join 4,200+ leaders getting practical insights every week — no fluff, just signal.

Join Free →

When it becomes a liability

The point at which an employee has acted on an answer nobody verified. Not necessarily to their detriment, and not necessarily in a way that produces a complaint. The exposure is the inability to reconstruct the exchange: somebody made a decision based on what a system told them, and no record exists of what it told them or where it got it. That is a records problem rather than a translation problem, and it is the one that becomes expensive.

The edge case that forces it

A policy change that lands across all populations at once, like a benefits restructure or a new leave arrangement. Every language set has to reflect it simultaneously, and the languages nobody reviews are the ones that will not. A single change event tends to surface every stale document you have, all at once, which is why the quarter following a major policy update is the highest-risk period in this whole arrangement.

Five Questions People Ask Before Setting This Up

"Can't we just use a translation tool to check?" You can, and it catches some things and systematically misses the most important one. A round-trip translation will reveal a garbled answer, a wrong register or an obvious mistranslation of a term. It will not reveal that the answer was drawn from a document that does not apply to the person asking, because that error survives translation perfectly intact. The round trip tests the language layer. The thing you need tested is the retrieval layer.

"How much reviewing is enough?" Far less than people assume, because you are hunting systematic errors rather than individual ones. Twenty answers per language per quarter, chosen to include the topics you know are jurisdiction-sensitive, will reliably surface a pattern. Reviewing everything is not more rigorous, it is just unsustainable, and a process that lapses after two months is worse than a small one that runs for two years.

"Who should do it?" Somebody who speaks the language and understands the policy, which is a smaller intersection than you would like. If you have to choose, pick the policy knowledge and get the language support, rather than picking a fluent speaker with no policy context. A bilingual colleague in another function can tell you an answer reads oddly. They usually cannot tell you it is wrong.

"Isn't this what the vendor should be doing?" Vendors test the model, not your documents, and the distinction is the whole problem. No vendor can tell you that your handbook lacks a section for one of your populations, because that is a fact about your organisation. Ask what they do test, and expect the answer to be about language model capability, which is useful but is not the same thing.

"What do we do with the findings?" Decide this before you start, because a review process with no route to a fix becomes a file nobody opens. Findings land in one of three buckets: the document was missing or wrong, the tool retrieved the wrong document, or the tool answered something it should have escalated. Each has a different owner, and naming those owners in advance is what turns a review into a process.

What You Are Actually Reviewing For

Three distinct failures hide under the phrase "the answer was wrong", and separating them is the difference between a useful review and a complaint log.

Translation failure. The source material was right and the rendering was wrong. A term of art translated literally, a register that reads as rude or absurdly formal, a negation dropped. These are the easiest to spot and the least consequential, because they usually look wrong to the reader, who then asks again or escalates. A reader who can tell something is off is a reader who is protected.

Retrieval failure. The right document exists and the tool used a different one. This is the dangerous category, and it is invisible to a reader who has no way of knowing which document informed the reply. The Spanish probation answer was a retrieval failure. It is also the failure most likely to be systematic, because whatever caused the tool to prefer the general document will keep causing it.

Coverage failure. The document the answer should have come from does not exist. Nothing is broken in the tool. The organisation simply never wrote the thing, and the tool did what it was designed to do, which is answer from the nearest available material. This is the category where the fix is not technical at all, and where the review process earns most of its value, because nothing else in your stack will surface it.

Failure What went wrong Visible to the reader Who fixes it
Translation Source was right, rendering was wrong Usually yes, so they ask again Whoever owns translation quality
Retrieval Right document existed, tool used another No, and this is the dangerous one Whoever administers the tool
Coverage The document that should answer it does not exist No. Nothing is broken Whoever owns the handbook
Out of scope Answer was correct but needed a person No, because it looked like an answer Whoever sets the configuration

A fourth is worth naming even though it is not strictly an error: an answer that was correct but should not have been given at all. Anything consequential enough that it depends on individual circumstance, and where the honest reply points at a person rather than a paragraph. A tool that answers these confidently is a problem even when the answer happens to be right, because it will not always be.

The Four Arrangements That Actually Work

A named internal reviewer per language

Somebody on the payroll who speaks the language, given an explicit slice of time for this. Not a volunteer, not a favour, a named responsibility with hours attached.

Right when you have a significant population in that language and therefore a decent chance of finding somebody suitable within it. It fails on two things: the reviewer is almost never in HR, so they need policy context supplied to them, and a single named person is a single point of failure during annual leave. Name a backup when you name the primary.

A reciprocal arrangement with the business

The country lead or a senior person in that market reviews a sample, and HR reciprocates with something they want, usually faster handling of their own queries or a say in the document priorities.

Right for mid-sized populations where a formal resource is hard to justify. It works because the market lead has a genuine interest in the answers being right. It fails when the relationship depends on goodwill that evaporates in a busy quarter, so put it in writing, with a cadence, however informal the spirit of it.

A paid external reviewer

A translation or localisation provider, briefed not to translate but to assess a sample of answers against your source documents for accuracy as well as language.

Right when nobody internal reads the language and the population is large enough that the risk justifies a cost. It fails if you brief it as a translation job, which is the usual mistake: a translator asked whether the Spanish is good will tell you the Spanish is good. The brief has to include the source document and the question of whether the answer follows from it.

Restricting what the tool answers at all

Not a review arrangement, but it belongs here because it is frequently the correct answer for a small population. Configure the tool to handle only the universal topics in that language, and to route everything jurisdiction-sensitive to a named human.

Right for the long tail, where a document set would be stale within a year and no realistic review process is affordable. It fails on expectation management if you do not tell people, because an employee who gets routed to a human without explanation experiences it as the tool being broken.

Building a Sample-Based Review That Survives Contact With a Busy Quarter

The process that works is small, specific and boring. Here is the whole of it.

Set the sample at twenty answers per language per quarter. Not a percentage, a flat number, because a percentage of a small population is too few to find anything and a percentage of a large one is unsustainable. Twenty is enough to surface a systematic pattern and small enough that it actually happens.

Stratify the sample deliberately rather than taking it at random. Half from your jurisdiction-sensitive topic list, which is where consequence lives. A quarter from the highest-volume questions, which is where scale lives. A quarter genuinely random, which is where surprises live. A purely random sample of twenty from a population whose questions are 80 per cent about expenses will tell you your expenses answer is fine.

Give the reviewer the question, the answer and the source document together. All three, every time. Without the source the reviewer cannot distinguish retrieval failure from coverage failure, which are the two findings that matter most and which have completely different owners. If your tool cannot export the source, that is a finding about your tool.

Ask two questions, scored separately. Is this fluent and appropriate in register. Does this answer follow correctly from the source document for this person's situation. Keeping the scores apart is the single most important design choice in the process, because collapsing them into one judgement is exactly how the nodding review happened.

Route findings to named owners the same week. Documents to whoever owns the handbook. Retrieval problems to whoever administers the tool. Escalation failures to whoever sets the configuration. A finding with no owner becomes a row in a spreadsheet and then becomes nothing.

Review the review once a year. Specifically, check whether the jurisdiction-sensitive topic list is still right, because policies change and new markets bring new divergences. A sample stratified against last year's risk list is a sample pointed in the wrong direction.

Building the Jurisdiction-Sensitive Topic List

Everything above stratifies against this list, so it has to exist before the review process does. It is also the artefact that gets referenced most and defined least, so here is how to produce one in an afternoon.

Start from your actual question history rather than from a template. Export six months of support questions, strip out anything about systems and access, and you will be left with policy questions grouped into perhaps thirty topics. That is your real surface area, and it is usually much smaller than the handbook's table of contents suggests.

Now sort those topics by one test: would the correct answer change if the person asking were employed in a different country. Not "might there be local nuance", which is true of everything and therefore useless as a filter. Would the answer change. Most topics fail that test and belong in the universal pile. The ones that pass are almost always clustered in the same handful of areas: notice and probation, statutory leave entitlements, sick pay and absence, parental arrangements, public holidays, working time, and anything touching the end of employment.

Then add a second column that nobody expects: who, internally, knows the answer for each market. Often the honest entry is nobody, or a provider, or an external adviser. That column is what makes the list usable, because a review finding on a topic where nobody internally knows the correct answer cannot be resolved by the reviewer and needs routing somewhere else entirely.

Keep the list short on purpose. Eight to twelve topics is a working list. Thirty is a handbook index with a new name, and a sample stratified against thirty topics is a sample stratified against nothing. If your list keeps growing past a dozen, the finding is that your policy set is fragmented rather than that your risk surface is wide, and that is a different project.

But the most useful property of this list has nothing to do with sampling. It tells you which questions your tool should be configured not to answer confidently. A topic on this list, asked by somebody in a market where the answer diverges and no local document exists, is a question where the right behaviour is to point at a person. Handing the list to whoever configures the tool is a five-minute conversation that prevents the single failure this whole article is about.

Revisit it once a year, and immediately after entering a new market. A new country does not just add a population, it frequently adds a divergence on a topic that was universal until that moment, which quietly invalidates the stratification you built the process on.

What a Launch Review Should Actually Cover

The team in the opening example did a launch review and it did not help, so it is worth saying what a useful one looks like, because the fix is small.

Test the same questions in every language rather than mostly in your working one. Thirty-one of forty in English tells you a great deal about English and almost nothing about the other four languages, and the distribution of that sample was the actual defect rather than its size.

Include at least three questions from the jurisdiction-sensitive list per language, and pick markets where you already know the answers differ. You are not testing whether the tool is fluent. You are testing whether it qualifies an answer when it should, and that behaviour only shows up on the questions where qualification is required.

Score fluency and accuracy separately, with the source document visible, and have the accuracy half done by somebody who knows the policy for that market. So if the honest answer is that nobody in the company knows the correct position for one of your markets, record that as a launch finding in its own right. It is more important than anything else the review will produce.

Finally, write down what the tool was configured to refuse, and test that too. A launch review that only checks the answers given, and never checks that the escalation path fires where it should, has tested half the system.

How to Choose: Five Questions Before You Talk to Any Vendor

Can you export the question, the answer and the source document together? Ask this in the first call and treat a no as close to disqualifying. Everything in this article depends on the third element, and plenty of tools store the conversation without storing what informed it.

Does the tool tell the reader when it is using general material? The behaviour you want is a hedge or a qualification when the retrieval is a loose match rather than a close one. A tool that answers with uniform confidence regardless of match quality transfers the entire risk to your review process.

Can you restrict topics by language or population? This is what makes the fourth arrangement above possible, and it is not universal. If you intend to let the tool answer only universal questions for your tail languages, confirm the configuration actually supports that rather than assuming.

What does the vendor test, and in which languages? Expect the honest answer to be model capability rather than answer accuracy, and do not penalise them for it. The value of asking is that it establishes clearly whose job the accuracy question is, which is yours.

How long are transcripts retained, and who can read them? A review process needs history, and a tool with a thirty-day retention window makes quarterly review impossible by arithmetic. Check the figure rather than assuming it is generous.

What the Tools Give You to Review With

Each product below was checked on its own site on 5 and 7 October 2026, and each is described on its own rather than alongside others, because mixing vendors and figures in one paragraph is how a price ends up attached to the wrong product.

Matram

Disclosure: Matram is owned by the same people who publish HROpsLab. It appears here because it competes in this category and is assessed against the same criteria as everything else on this page, with its limitations stated in the same detail.

Best for: teams whose review problem is caused by breadth, because it publishes 95+ languages and prices flat at $29, $69 and $199 a month with seats unlimited, so the cost of the arrangement does not rise with the number of people involved in reviewing it.

Why it helps here: it answers from your own documents rather than from general knowledge, which is what makes source attribution meaningful in the first place. A document-trained tool can tell you which of your materials produced an answer. A general assistant cannot, because the answer did not come from a document at all. The 30-day trial with no card required also means the sample-based review described above can be run as an evaluation rather than after purchase, which is the right order.

Where it struggles: no free tier, so a tail-language pilot cannot be parked indefinitely. It answers rather than actioning tickets, so it is not a workflow platform. And it does not solve the underlying problem any better than its competitors do, because the coverage failure category is a fact about your handbook and no vendor can fix that.

SiteGPT

Also document-trained, and also publishes 95+ languages, so on reach the two are level and the choice rests elsewhere. Priced at $468 and $948 billed yearly, with the monthly-equivalent figures applying only on that annual commitment. Starter covers one chatbot and a capped page count, which matters if your plan was a separate deployment per population for review isolation.

Chatling

Has a free tier, which makes it the cheapest way to run a review experiment before committing. Its published language figure is inconsistent, reading 80+ in one homepage section and over 85 in another, which is a small discrepancy and a useful reminder of how lightly these numbers are held across the category.

Botsonic

Publishes 50+ languages, a narrower and plainly stated figure. For review purposes a smaller number you can trust is more useful than a larger one you cannot, so check your specific language list against it rather than assuming the tail is covered.

Crisp

Priced per workspace at Free, $45, $95 and $295 a month, so adding reviewers and agents does not move the bill, which is a genuine advantage when a review process means more people with access. It publishes no language count, so the coverage question has to be answered in a call.

Zendesk

Suite Team is $55 per agent per month with the Copilot add-on at a further $50 per agent per month. The per-seat basis is the awkward part for review specifically, because every additional reviewer you give access to is another licence. It publishes no language figure. Its pricing page also geo-redirects, so confirm the currency.

Freshservice

Tiered per agent at $19, $49 and $99 per month, with Freddy AI priced separately at $29 per agent per month. Same structural issue for review access as any per-seat product. No published language count.

Three vendors publish no price at all. Leena AI and Moveworks are demo-only, with Moveworks lacking a working pricing page. Tidio has replaced plan pricing with a usage calculator, so the figures on its page are conversation counts rather than money.

The Comparison

Tool Document-trained Language count published Pricing basis Cost of adding a reviewer
Matram Yes 95+ Flat, seats unlimited None
SiteGPT Yes 95+ Annual per plan None within plan
Chatling Yes Inconsistent, 80+ and over 85 Per plan, free tier None within plan
Botsonic Yes 50+ Per plan None within plan
Crisp Partially None published Per workspace None
Zendesk Partially None published Per agent plus AI add-on Two licence lines
Freshservice Partially None published Per agent plus AI add-on Two licence lines

The Decision Table

Situation Scale Setup Primary Pain Recommended Starting Point
HR team reads every supported language Any Ordinary QA None specific to language Standard sample review. Skip this arrangement
Significant population, speakers available internally 200 to 1,000 Named internal reviewer plus backup Reviewer lacks policy context Name the reviewer, supply the source documents, allocate hours
Mid-sized population, no formal resource 200 to 1,000 Reciprocal arrangement with market lead Depends on goodwill in a busy quarter Put the cadence in writing, however informal the spirit
Large population, nobody internal reads it 500 plus Paid external reviewer Briefed as translation, so accuracy goes unchecked Brief against source documents, not language quality
Long tail of two and three person populations Any Restrict topics, route the rest Document sets go stale unread Limit the tool to universal topics. Tell people plainly
Tool cannot export source documents Any Any Retrieval and coverage failures are indistinguishable Make source export a shortlisting requirement before buying
Major policy change just landed Any Existing process Every stale language set surfaces at once Run an off-cycle review of all languages that quarter

When the Reviewer and HR Disagree

This happens more than the tidy version of the process admits, and it has a specific shape. A reviewer in-market says an answer is wrong. The HR team reads the source document and says it follows from the handbook. Both are correct, and what they have actually found is a coverage failure that looks like a dispute.

Resolve it by asking one question: is the handbook right for this market. Not is the answer faithful to the handbook, which is the question HR is answering, and not does the answer match local practice, which is the question the reviewer is answering. Those two can both be yes while the handbook is wrong, and the review process is the only mechanism likely to surface that.

So treat a disagreement as the most valuable output the process produces rather than as friction to be settled. Escalate it to whoever can confirm the local position, which is frequently an external adviser rather than anybody internal, and record the outcome in the handbook rather than only in the review log. A disagreement resolved in a transcript thread fixes one answer. The same disagreement resolved in the source document fixes every future answer, which is the entire point of having a document-trained tool.

The Signals That Tell You Quality Has Slipped

You will not get complaints, so watch for these instead.

Questions per head falling in one population. The most reliable signal in this whole article. A group filing half as many questions per person as the company average has not become better informed. They have stopped asking, usually because the answers were not useful and escalating in a second language is effortful.

Repeat questions from the same person within a short window. Somebody asking the same thing three different ways is somebody who did not believe or did not understand the first answer. In a single-language setup this reads as a findability problem. In a multilingual one it is frequently an accuracy problem.

Escalations that bypass the tool entirely. Managers fielding policy questions directly for their team is the clearest evidence that the tool is not trusted in that language, and it is invisible in tool analytics because the interaction never happened there. Ask managers directly, periodically.

A document set with no edits since launch. Check the modification dates. A second-language document set that has not changed since go-live, through two policy updates, is stale by definition and nobody noticed, which tells you the review process is not running whatever the calendar says.

What Getting This Wrong Costs

The direct cost is small and the indirect cost is the whole point. Setting up what this article describes is a named person, twenty answers a quarter and an hour of somebody's time, which at any realistic rate is a trivial line next to the licence.

The second cost is the decision made on a wrong answer. Most of these are minor and recoverable, and the occasional one is not. What makes it worse than an ordinary mistake is the absence of a record: when somebody asks what the employee was told, the honest answer is often that nobody knows, because the transcript stored the reply and not its basis. That turns a correctable error into an unresolvable dispute.

The third cost is the one that compounds. A population that learns the system is confident and unreliable disengages from it permanently, and they do not come back when you fix it, because nobody announces a fix to people who stopped asking. You lose the channel, and you lose it quietly, and your metrics improve while it happens.

So the question worth asking before launch is not whether the answers are good. It is whether anybody will be able to tell, six months from now, that one of them was wrong.

When You're Ready to Move Beyond Spot-Checking

Almost every team starts with spot-checking, and it is a perfectly reasonable place to start. Somebody reads a few answers after launch, it looks fine, and attention moves on. This is not negligence; it is what a launch looks like when the launch goes well.

What makes it stop working is time rather than scale. Spot-checking catches the errors that exist on day one and is structurally blind to the ones that arrive later, which are the majority, because they arrive with policy changes and document drift rather than with the implementation. The case for a sample-based process is not that it is more rigorous on launch day. It is that it still exists in month fourteen, which spot-checking never does.

The sequence that works: build the jurisdiction-sensitive topic list first, because everything else stratifies against it. Name the reviewers and the backups. Confirm your tool can export the source document alongside the answer, and if it cannot, treat that as the finding it is. Then run twenty answers a quarter and route the results to people by name. It is a small process, and it is the only thing standing between a fluent system and a trusted one.


Frequently Asked Questions

How do you review HR answers in a language nobody in HR speaks?

By separating the two questions you are actually asking and getting different people to answer them. Fluency and register can be assessed by any competent speaker of the language, including a colleague outside HR or an external provider. Whether the answer follows correctly from your own policy document requires the source document in front of the reviewer, which is why exporting the question, the answer and the source together is the single capability this whole process depends on.

Is a round-trip translation enough to check an answer?

It catches some things and misses the most important one. Translating the answer back into English will reveal garbled language, a wrong register or an obviously mistranslated term, all of which are worth catching. What it cannot reveal is that the answer was drawn from a document that does not apply to the person who asked, because that error translates perfectly and arrives sounding entirely reasonable in both languages.

How many answers should we review?

Twenty per language per quarter is a good default, chosen as a flat number rather than a percentage so that it stays meaningful for small populations and sustainable for large ones. Stratify it rather than sampling at random: half from the topics you know differ by jurisdiction, a quarter from your highest-volume questions, and a quarter genuinely random. You are hunting systematic patterns, and twenty stratified answers will surface one.

Who should do the reviewing?

Somebody who has both the language and the policy context, and where you cannot find that combination, favour policy knowledge and supply language help rather than the reverse. A fluent colleague with no policy background can tell you an answer reads oddly but usually cannot tell you it is wrong, which is the judgement you actually need. Name a backup at the same time as the primary, because a single reviewer means an entire language goes unchecked during annual leave.

What is the difference between a retrieval failure and a coverage failure?

A retrieval failure means the right document existed and the tool used a different one, so the fix is in the tool's configuration or your document structure. A coverage failure means the document that should have answered the question does not exist, so the tool answered from the nearest available material and nothing is technically broken. They look identical in the transcript and have completely different owners, which is why the reviewer needs to see the source document alongside the answer.

Should the vendor be testing answer accuracy?

Vendors test the capability of the underlying language model, which is useful and is not the same thing as testing whether your answers are right. No vendor can tell you that your handbook is missing a section for one of your populations, because that is a fact about your organisation rather than about their product. Asking the question is still worth doing, because it establishes clearly that the accuracy question belongs to you.

How do we know quality has slipped if nobody complains?

Watch questions per head by population rather than waiting for complaints, because the usual response to an untrustworthy answer is to stop asking rather than to escalate. A group filing noticeably fewer questions per person than the company average, repeat questions from the same individual in a short window, and managers informally fielding policy questions for their team are the three reliable signals. Also check the modification dates on your second-language documents, since a set unchanged since launch through two policy updates is stale regardless of what the review calendar says.

HROpsLab takes no vendor money and publishes no paid placements, which is why this page says "none published" three times rather than estimating.

Share on X Share on LinkedIn

What to do next?

Explore More Articles

Dig deeper into HR Ops strategy, tools, and workflows built for real teams.

Browse the blog →
Join the HROpsLab Community

Connect with People Ops practitioners sharing real workflows, tools, and challenges.

Join now →