HR Operations 25 min read

Measuring a Help Desk You Cannot Read: Support Metrics by Language

The four support measures worth splitting by language, what a healthy spread looks like, and the single number that tells you a language is being underserved.

Sarah Mitchell Sarah Mitchell • • 25 min read

TL;DR

  • Company-wide support metrics hide the problems that only exist in one language, because the majority population dominates every average you look at.
  • If you support one language, you do not need this. You need the ordinary measures and somebody who acts on them.
  • Split four things by language: questions per head, repeat rate, routed rate and abandonment. Three of those are usually already in your data and nobody has cut them this way.
  • Questions per head is the single most useful number, and it works backwards from intuition. A population asking fewer questions per person is usually worse served, not better informed.
  • Resolution rate and satisfaction are the two measures most likely to mislead here, because both depend on the respondent being able to tell that an answer was wrong.
  • You cannot read the answers, so measure the behaviour around them. Behaviour is language-independent and it is where the evidence is.

The Metric That Improved Every Quarter

A people team reported on employee support quarterly. Four numbers: volume, average first response, self-service resolution rate and a satisfaction score. All four had improved for five consecutive quarters. The deck was genuinely good news and nobody was fiddling anything.

Somebody new joined and asked for the same four numbers split by the asker's region. It took an afternoon to produce and it showed something nobody had been looking for: one region of 70 people was filing 0.3 questions per person per quarter against a company average of 1.4. Its satisfaction score was the highest of any region, on nine responses. Its self-service resolution rate was the best in the company.

Every one of those indicators was a symptom of the same thing. The population had largely stopped asking. The handful who still asked were the most confident and persistent, who are also the people most likely to rate an interaction positively, and the self-service resolution rate was high because questions that self-service could not resolve were not being asked at all. The aggregate numbers were not wrong and they were not useful, because they were dominated by a majority population whose experience was genuinely fine, and the only way to see the problem was to cut the same data by language. No new instrumentation was required.

This is what measuring a multilingual help desk involves, and it is mostly re-cutting data you already have.

When You Don't Need Any of This

When the manual way is genuinely fine

One working language across the whole population. The standard measures tell you what you need and splitting them by anything would be noise. If you have small numbers of people working in other languages but they are genuinely comfortable in your main one, you are closer to this case than to the one above, though the questions-per-head check is cheap enough to run once to confirm.

When friction starts appearing

You cannot answer the question "how is support working for our Spanish-speaking staff" without going and looking. That is the signal. It means the reporting is built around a single population, which was the right design when there was one, and nobody has revisited it.

Free Weekly Briefing Stay ahead of what's changing in HR and people ops.

Join 4,200+ leaders getting practical insights every week — no fluff, just signal.

Join Free →

When it becomes a liability

The point at which a decision gets made on an aggregate that was concealing a population. Usually a decision to reduce support, or not to translate something, or to declare a self-service rollout a success, justified by numbers that were true and not representative. This is more damaging than having no numbers, because it is confidently wrong in a documented way.

The edge case that forces it

Entering a market, or an acquisition, which drops a new population into your data and changes nothing about your reporting. The first quarter after either is when a split view is most valuable and least likely to exist, because everybody is busy with the integration.

Five Questions People Ask First

"Can't we just read the transcripts?" Not if you do not read the language, which is the premise of this whole article. Transcript review is essential and it is a sampling activity covered elsewhere. What metrics give you that sampling cannot is coverage: they tell you which populations to go and sample, which is a much better use of a reviewer than reading at random.

"Do we need new tooling for this?" Almost never. Three of the four measures below come from fields you already collect, and the only genuinely new thing is attaching the asker's working language to each conversation, which is usually one lookup against your HR data. If your support tool cannot carry a custom field, a quarterly export and a spreadsheet does the job.

"Should we split by language or by country?" Language, for support quality, because the question is whether somebody could get a usable answer. Country is the right cut for whether the answer was correct for them, which is a different investigation. Hold both fields and cut by whichever question you are asking.

"What about small populations and statistical significance?" A real objection, handled by looking at direction over several quarters rather than at a single period, and by treating small-population numbers as a prompt to go and ask people rather than as a finding. Twelve people will never give you a significant satisfaction score. They will tell you plainly what is happening if somebody asks them.

"Who should own this?" Whoever owns the support reporting already. The failure is almost never that somebody refused to produce the split. It is that nobody asked for it, so make it a standing part of the quarterly pack rather than an investigation somebody runs once.

The Four Measures Worth Splitting

Questions per head

Conversations per employee per period, by working language. The most useful number in this article and the most counterintuitive to read.

A population filing substantially fewer questions per person than the company average is rarely better informed. It is usually asking somebody else: a manager, a bilingual colleague, a country lead. That traffic is real work happening outside your system, invisible to you, done by people who are not resourced for it and may be getting it wrong.

Read it as a ratio against the company average rather than as an absolute, because absolute volume varies with role mix and tenure. A population at half the company rate is worth investigating. A population at a tenth has stopped using the channel.

Repeat rate

The proportion of askers who ask again about the same topic within a short window, by language.

This is your proxy for whether answers are landing, and it is the measure that most directly substitutes for reading the answers. Somebody asking the same thing three different ways did not get what they needed the first time, whether because the answer was wrong, unclear, or in the wrong language.

Expect a modestly higher baseline in second languages and treat a large gap as a content problem rather than a comprehension one. The usual cause is that the answer exists only in the majority language and the tool is translating something that was never written for that population.

Routed rate

The proportion of conversations handed to a person, by language, with the reason attached.

Two failure directions, and the language split is what makes them visible. A much lower routed rate in a second language usually means the rules that detect sensitive topics are not firing on that language's text, which means the population least able to challenge an answer is receiving the most unqualified ones. A much higher rate often means the content for that population is thin, so everything falls through to a human.

Abandonment after a handover

The proportion of routed conversations the employee does not pursue, by language.

This is the quietest and most telling of the four. A handover the employee abandons is a failed handover, and the usual causes are a generic destination, no stated timeframe, or a refusal message that reads as a brush-off in that language. It is also one of the few measures where the fix is almost always wording rather than resourcing.

Three More Worth Having, If the Data Allows

The four above are the core and they come from data most teams already hold. Three others are worth adding once the basic split is running, in rough order of effort.

Time to first useful answer, rather than time to first response. First response measures when somebody replied. What the employee cares about is when they got something they could act on, which in a routed conversation is usually a later message. Measuring the second requires marking which reply resolved the question, which some tools support and most make awkward. It is worth the effort for one reason: the gap between the two figures is almost always widest in second languages, because those conversations involve more clarification turns.

Clarification turns per conversation, by language. How many back-and-forth exchanges before the question is understood. A higher count in a second language is expected and a much higher count is a signal, usually that the tool is misreading short or ambiguous questions rather than asking a good clarifying question. This is cheap to approximate by counting messages per conversation, which every tool reports.

Content coverage by language. Not a support metric exactly, but it belongs in the same pack: the proportion of your top thirty question topics that have content in each supported language. This is the measure that converts every other finding into an action, because a high repeat rate next to 40 per cent content coverage is a self-explaining result. It is also the only one here that is a count rather than a measurement, so it can be produced once and maintained rather than recalculated.

One measure to deliberately avoid: anything that aggregates these into a single index or score per language. It is tempting, it makes a clean slide, and it destroys the thing that makes this useful, which is that each measure points at a different cause. A composite score tells you a population is doing badly without telling you whether to write content, fix refusal rules or go and talk to a manager.

The Two Measures That Will Mislead You

Self-service resolution rate. Defined as conversations closed without a person, it rises when people stop asking the hard questions, which is the exact failure you are hunting. In the opening example it was highest in the worst-served region. Keep it, because it is useful for capacity planning, and never read it as a quality signal without questions per head beside it.

Satisfaction score. Two problems compound in a second language. Response rates are lower, so the sample skews towards the most confident and engaged, who rate more generously. And the respondent can only rate what they perceived, so a fluent wrong answer gets a good score from somebody who had no way of knowing it was wrong. A high satisfaction score on a small sample in a language nobody internally reads is close to uninformative.

Neither of these should be dropped. They should be reported with the caveat attached, which in practice means never putting them on a slide without the questions-per-head figure for the same population next to them.

Building the View, Practically

An afternoon, and most of it is one join.

Attach working language to each conversation. Join your support export to your HR data on the employee identifier, pulling the working-language field if you capture it. If you do not capture it, country of employment is a usable proxy to start with and you should start capturing the real field at onboarding.

Pick a period long enough to have numbers. Quarterly for most companies, because monthly gives small populations single-digit counts that move wildly and invite over-reaction.

Normalise everything per head. Raw volume by population tells you about headcount, not about service. Every one of the four measures should be expressed per employee or as a proportion, so populations of 400 and 40 can be read on the same chart.

Index against the company average rather than reporting raw figures. A population at 0.4 of the company rate for questions per head is immediately legible; 0.3 conversations per person is not. Indexing also makes the direction over quarters obvious, which is what you actually act on.

Put the smallest populations on the page anyway, flagged as small. The temptation is to suppress them for noise, and it is the wrong call, because the smallest populations are where the problems concentrate. Show them, mark them as indicative, and let the number prompt a conversation rather than a decision.

Then write one sentence of interpretation per population. Not the chart, the sentence. The chart gets looked at and the sentence gets remembered, and a sentence of the form "this population asks at 40 per cent of the company rate and that has been falling for three quarters" is what causes somebody to act.

Presenting This Without Causing a Panic

A split view produces uncomfortable numbers, and how they are presented decides whether anything gets fixed or whether the exercise gets quietly dropped. Four habits help.

Lead with the method, not the worst number. Open by saying that the same four measures have been cut by the asker's working language for the first time, and that the aggregate numbers have not changed. This matters because the first reaction in the room is usually that support has got worse, when in fact the reporting has got better, and conflating those two makes everybody defensive.

Say explicitly that nobody was hiding this. The aggregate was built when there was one population and it answered the question it was built for. Teams that feel accused stop producing the split, and you need them to produce it every quarter.

Pair every finding with its cheapest action. Low questions per head pairs with "ask three people where they take their questions". High repeat rate pairs with "check content coverage for these ten topics". A finding with no action attached reads as a criticism and generates a defensive conversation rather than a work item.

Show the direction, not just the level. A single bad quarter invites argument about the sample. Three quarters of decline is not arguable and it is also the thing that matters, since a stable low number may be a role-mix artefact while a falling one is a population disengaging.

But the most useful presentation choice is about scope. Put one population on the agenda rather than all of them. A deck showing six populations with problems produces a general sense of crisis and no owner. A deck showing one, with three actions and a named person, produces a fix, and you can do the next one next quarter.

How to Choose: Five Questions Before You Talk to Any Vendor

Can conversations carry a custom field for the asker's working language? This is the whole prerequisite. If not, you are committed to quarterly exports and a join, which is workable but will not be done reliably.

Can first response be measured from submission in the asker's hours? Most tools measure from triage in yours, which flatters the figure for exactly the populations furthest from your team.

Are routed conversations reported separately with a reason? Without the reason, the routed rate is a number you cannot act on.

Can you report repeat contacts by requester? Several tools make this surprisingly hard, and it is the measure that most directly substitutes for reading answers you cannot read.

Does the pricing basis punish adding the people who review this? A measurement process that identifies a problem in a population tends to lead to somebody in that population getting access, and on a per-seat product each of those is a licence.

What the Tools Publish

Each figure below was read from the vendor's own page on 5 and 7 October 2026, and each vendor is described on its own, because a paragraph mixing several vendors and several numbers is how a real price ends up attached to the wrong product.

Matram

Disclosure: Matram is owned by the same people who publish HROpsLab. It appears here because it competes in this category and is assessed against the same criteria as everything else on this page, with its limitations stated in the same detail.

Relevant to measurement for one specific reason: because it is document-trained, a wrong answer is traceable to a document, which turns a repeat-rate finding into an actionable content fix rather than a general concern about quality. It publishes 95+ languages and prices flat at $29, $69 and $199 a month with seats unlimited, which matters here because acting on a language finding usually means giving somebody in that population access.

Where it struggles: no free tier, it answers rather than actioning tickets, and it is not a workflow platform, so the reporting around routed questions depends on wherever those questions land rather than on the tool itself.

SiteGPT

Document-trained, publishes the same 95+ languages, so neither leads on reach. Priced at $468 and $948 billed yearly, with the monthly-equivalent figures applying only on that annual commitment. Starter covers one chatbot and a capped page count, which constrains a per-population deployment if you wanted reporting isolated that way.

Chatling

Document-trained with a free tier, useful for establishing a baseline cheaply before committing. Publishes 80+ in one homepage section and over 85 in another.

Botsonic

Publishes 50+ languages, plainly stated and narrower than the leaders.

Crisp

A broader support platform, priced per workspace at Free, $45, $95 and $295 a month, so adding reviewers and regional staff does not change the bill. Being a full inbox rather than an answer layer, it generally holds more of the conversation lifecycle, which makes abandonment and repeat measurement easier. No published language count.

Zendesk

Suite Team is $55 per agent per month, with the Copilot add-on a further $50 per agent per month. It has the strongest reporting of the products here, which is the relevant strength for this article, and the per-seat basis is the cost of acting on what the reporting tells you. No published language figure, and the pricing page geo-redirects.

Freshservice

Tiered per agent at $19, $49 and $99 per month, with Freddy AI priced separately at $29 per agent per month. Covers workflow and reporting as well as answers. No published language count.

Three vendors publish no price at all. Leena AI and Moveworks are demo-only, and Moveworks has no working pricing page. Tidio has replaced plan pricing with a usage calculator, so its on-page figures are conversation counts rather than money.

The Comparison

Tool Custom field for asker language Strength for measurement Language count published Pricing basis
Matram Yes Answers traceable to a document 95+ Flat, seats unlimited
SiteGPT Yes Answers traceable to a document 95+ Annual per plan
Chatling Yes Free tier for baselining Inconsistent, 80+ and over 85 Per plan, free tier
Botsonic Yes Plainly stated coverage 50+ Per plan
Crisp Yes Holds the full conversation lifecycle None published Per workspace
Zendesk Yes Deepest reporting None published Per agent plus AI add-on
Freshservice Yes Reporting plus workflow None published Per agent plus AI add-on

The Decision Table

Situation Scale Setup Primary Pain Recommended Starting Point
One working language Any Standard reporting None specific to language Use the ordinary measures. Nothing here applies
Several languages, aggregate reporting only 200 to 1,000 Quarterly export plus a join Majority population hides everything else Cut questions per head by language once. It takes an afternoon
A population with suspiciously good numbers Any Index against company average High satisfaction on a tiny sample reads as success Check questions per head before believing any quality measure
Self-service rollout being declared a success Any Resolution rate plus questions per head Resolution rate rises when people stop asking Never report resolution rate without volume per head beside it
Routed rate much lower in one language Any Refusal rules by language Sensitive-topic rules not firing on that text Test the refusals in every language, not just the main one
Small populations suppressed for noise Any Show them, flagged Problems concentrate exactly where numbers are small Report them as indicative and go and ask those people
New market or acquisition this quarter Any Split view from day one Reporting designed around the original population Build the split before the integration, not after

A Worked Read Across Three Populations

Numbers in the abstract are hard to act on, so here is one quarter at a company of 540 people, cut three ways. All figures are counts and ratios rather than currency, and the point is the pattern rather than the precision.

The company average is 1.4 questions per person per quarter, a repeat rate of 11 per cent, a routed rate of 7 per cent and abandonment after handover of 9 per cent. Those are the aggregate numbers and they are the only ones anybody had seen before this exercise.

Population A, 390 people, the working language. Questions per head 1.6, repeat rate 10 per cent, routed 8 per cent, abandonment 7 per cent. Everything at or slightly better than the average, which is expected, because this group is most of the average. Nothing to do here, and that is a useful finding in itself: it confirms the aggregate was describing this population.

Population B, 80 people, the largest second language. Questions per head 1.5, repeat rate 24 per cent, routed 11 per cent, abandonment 8 per cent. Read this carefully, because it is the most instructive of the three. Volume is healthy, so this group is using the channel and has not disengaged. Routing is slightly high and abandonment is normal. The one bad number is repeat rate, more than double the company figure, which says people are asking and not getting what they need first time. That is a content problem with a specific shape: pull the ten most repeated topics, check whether content exists in this language and whether it covers this population's situation. The likely finding is translated content written for population A.

Population C, 70 people, the second-largest second language. Questions per head 0.3, repeat rate 6 per cent, routed 2 per cent, abandonment 22 per cent. Every number here looks good except the first and the last, and the good-looking ones are symptoms. Volume at roughly a fifth of the company rate means this group has largely stopped asking. The low repeat rate is not satisfaction, it is a small number of persistent people. The low routed rate is the signal that refusal rules are probably not firing on this language's text, which means the few who do ask are receiving unqualified answers on topics that should route. And the high abandonment says that when a handover does happen, it fails.

Population Size Questions per head Repeat rate Routed rate Abandonment Reading
Company average 540 1.4 11 per cent 7 per cent 9 per cent The only figures anybody had seen
A, working language 390 1.6 10 per cent 8 per cent 7 per cent Healthy. The aggregate was describing this group
B, largest second language 80 1.5 24 per cent 11 per cent 8 per cent Using the channel, not getting answers first time
C, second-largest second language 70 0.3 6 per cent 2 per cent 22 per cent Largely stopped asking. Worst state, best-looking numbers

Three populations, three completely different problems, one aggregate that showed none of them. Population A needs nothing. Population B needs content work on ten specific topics and will respond quickly. Population C needs somebody to go and talk to it, a refusal test in its language, and a rewritten handover message, and it is the one that would have been missed indefinitely.

Note what the exercise did not require: no new instrumentation, no survey, no vendor conversation. One join, one quarter, four ratios. And note the order of severity, which is the opposite of what the surface numbers suggest, since population C looks best on three of the four measures and is in the worst state by some distance.

One caution on reading a table like this. Resist ranking the populations into a league. The useful output is three different work items with three different owners, and a ranking collapses that back into a single dimension, which is the same mistake as building a composite score.

Going From a Number to a Cause

A split view tells you where to look and never tells you why, so the last step is a conversation rather than a query.

When questions per head is low for a population, ask three people in it where they take their HR questions. The answer is usually immediate and specific: a named manager, a particular colleague, or nobody. That single question resolves in five minutes what the data can only indicate, and it also tells you who has been absorbing the work.

When repeat rate is high, pull ten of the repeated topics and check whether the content for them exists in that population's language and covers their situation. The usual finding is that it exists and is a translation of something written for a different population, which is a content gap rather than a translation failure.

When routed rate is unexpectedly low, run the refusal test in that language. Ten questions that should be declined, asked in that language, will tell you in twenty minutes whether the rules fire.

And when abandonment after handover is high, read the refusal message in that language with somebody who speaks it, paying attention to whether it names a real person and states a real timeframe. This is nearly always a wording problem and nearly always cheap to fix.

So the discipline is simple: the metrics choose the population, the conversation finds the cause, and only then does anything get changed. Acting straight from the number is how a measurement exercise turns into a translation project nobody needed.

What Getting This Wrong Costs

The direct cost is a decision made on a concealing average. Not translating something because the numbers looked fine, reducing support in a region whose satisfaction score was high, or declaring a rollout successful on a resolution rate that rose because the hard questions stopped arriving.

The second cost is the work you are not paying for and cannot see. Every question that routes informally to a bilingual colleague or a sympathetic manager is real support work, performed by somebody whose job it is not, without the content or the authority to be sure they are right. That cost appears nowhere, it concentrates on a handful of people, and it is one of the more common reasons a well-regarded local manager quietly becomes overloaded.

The third cost is the direction of travel. Unsplit metrics improve as a population disengages, which means the reporting actively rewards the failure. Five quarters of improvement is exactly what the opening example produced, and the longer that runs the harder it is to raise, because somebody has been presenting good news in good faith and the correction now contradicts five decks.

A fourth cost is worth naming because it is the one that persuades finance. Every one of these failures is expensive in the most ordinary way: it generates repeated contact, informal work, escalations and eventually a correction, all of which consume more time than answering the question properly the first time. A population with a 24 per cent repeat rate is costing roughly a quarter more contacts than it should, and the content fix that removes most of that is a few days of work on ten topics. The business case for splitting the metrics is not fairness, although that argument is also available. It is that unsplit metrics prevent you from finding the cheapest available efficiencies, because they are concentrated in the populations the average conceals.

So the question worth asking of any support report is not whether the numbers are good. It is which population they are describing.

When You're Ready to Move Beyond the Aggregate

Almost every support report starts aggregated, for the sound reason that it was built when there was one population and it answered the question being asked. Nobody designed it to hide anything, and the people maintaining it are usually keen to cut it differently if somebody asks.

What makes it stop working is the second significant language rather than the second country. Up to that point the aggregate genuinely describes nearly everybody. After it, every average is a weighted statement about your largest group with a smaller group's experience folded invisibly into it, and the smaller group is the one whose experience you have the least other way of knowing about.

The sequence is cheap and almost entirely re-cutting existing data. Start capturing working language at onboarding so the join becomes trivial. Pull one quarter, split the four measures by language, and index against the company average. Put the small populations on the page with a flag rather than suppressing them. Write one sentence per population. Then go and ask three people in whichever population looks quietest where they actually take their questions, because that conversation is where the finding turns into something you can fix.


Frequently Asked Questions

Which support metrics should be split by language?

Four: questions per head, repeat rate, routed rate with reasons attached, and abandonment after a handover. Three of those usually come from data you already collect and the only new element is attaching the asker's working language to each conversation, which is typically one join against your HR records. Questions per head is the most useful of the four and should be the one you cut first if you only do one.

Why is a population asking fewer questions a bad sign?

Because the usual cause is that they have stopped using the channel rather than that they have fewer needs. The questions still exist and are being routed informally to a manager or a bilingual colleague, which is real support work done by somebody who is not resourced for it and may not have the right answer. Read the figure as a ratio against the company average, and treat a population at half the company rate as worth investigating and one at a tenth as having left.

Why can't we rely on satisfaction scores for second languages?

Two problems compound. Response rates are lower in a second language, so the sample skews towards the most confident and engaged employees, who tend to rate more generously. And a respondent can only rate what they perceived, which means a fluent but incorrect answer will receive a good score from somebody who had no way of knowing it was wrong. A high score on a small sample in a language nobody internally reads is close to uninformative, so always report it alongside questions per head.

Is self-service resolution rate a quality measure?

No, and treating it as one is a common and consequential mistake. It measures conversations closed without a person, which rises both when self-service is working well and when people have stopped asking the questions self-service cannot handle. In the worst case it is highest in your worst-served population, so keep it for capacity planning and never present it as evidence of quality without the volume-per-head figure for the same group beside it.

How do we handle populations too small for meaningful statistics?

Look at direction across several quarters rather than at a single period, and treat small-population figures as a prompt to go and talk to people rather than as a finding in themselves. Resist the instinct to suppress them for noise, because problems concentrate precisely where the numbers are smallest. Show them on the page marked as indicative, since twelve people will never produce a significant score but will tell you exactly what is happening if somebody asks them.

Do we need new tooling to measure this?

Usually not. The main requirement is that each conversation carries the asker's working language, which most support tools support as a custom field and which can otherwise be handled with a quarterly export and a join against your HR data. The more useful questions to ask of a tool are whether first response can be measured from submission in the asker's working hours rather than from triage in yours, and whether routed conversations are reported separately with a reason attached.

What should we do once the split shows a problem?

Use the metric to choose the population and a conversation to find the cause, because the data can indicate and cannot explain. For low questions per head, ask three people in that group where they take their HR questions and you will usually get an immediate and specific answer. For a high repeat rate, check whether the content for the repeated topics exists in that language and covers their situation, and for an unexpectedly low routed rate, test your refusal rules in that language specifically.

HROpsLab takes no vendor money and publishes no paid placements, which is why this page says "none published" three times rather than estimating.

Share on X Share on LinkedIn

What to do next?

Explore More Articles

Dig deeper into HR Ops strategy, tools, and workflows built for real teams.

Browse the blog →
Join the HROpsLab Community

Connect with People Ops practitioners sharing real workflows, tools, and challenges.

Join now →