Measuring Performance Without Distorting It

Every measure of an individual changes the work, because people optimise for what is counted and the counting is never the whole job. Six kinds of measure reviewed against the behaviour each produces, and how to assess work where the output is judgement.

Rachel Kim Rachel Kim 24 min read
Measuring Performance Without Distorting It

TL;DR

  • The core decision: which distortion you're willing to accept, because every measure of an individual's work produces one.
  • When doing nothing is right: when a manager's judgement is trusted, decisions are explainable, and nobody has asked for numbers.
  • What has to be true: you can say what behaviour each measure will encourage, before it's introduced rather than after.
  • How the options split: by what they count, and by how much of the job that counting represents.
  • Decision rule: pair anything countable with a judgement about the part that isn't countable, or the countable part becomes the job.
  • Outcome to expect: fewer measures, chosen deliberately, with their side effects named in advance.

The Request for Objectivity

Somebody asks for performance to be more objective. The reasoning is good: the current assessments rest on manager judgement, which varies between managers, favours visible people, and is difficult to defend when challenged. Numbers would be fairer.

So a list gets assembled of things that can be counted. In most jobs there are a few: volume of output, response times, error rates, whatever the systems happen to record. The list looks reasonable, it's certainly more objective than an opinion, and it becomes the basis of assessment.

Then the behaviour changes, and it changes fast. People start doing more of whatever is counted and less of whatever isn't. The support person whose response time is measured gets back to everybody quickly and resolves less. The developer measured on tickets closed picks the small ones. Nobody is cheating. They're doing exactly what the measure asked for, and the measure asked for something narrower than the job.

The real problem isn't that measurement is wrong. It's that every measure of an individual's work changes the work, because people optimise for what's counted and the counting is never the whole job. That's not an argument for measuring nothing. It's an argument for choosing measures whose distortion you can live with, keeping them few, and naming the side effects before you introduce them rather than discovering them a quarter later.

When You Genuinely Do Not Need to Act Yet

Your current setup is genuinely fine. Managers know the work well enough to assess it, decisions can be explained, and nobody has raised fairness as a concern. Judgement that's trusted and explainable is a perfectly good basis for assessment, and adding metrics to a working system usually costs more than it returns.

Friction is starting to show. Two managers have assessed similar work differently, or somebody has questioned how a decision was reached, or you can't answer a straightforward question about how a team is performing. That's a consistency problem and it may have a measurement answer, though it may equally have a definitions answer.

Free Weekly Briefing Stay ahead of what's changing in HR and people ops.

Join 4,200+ leaders getting practical insights every week — no fluff, just signal.

Join Free →

It has become a real cost. Decisions are being contested, or quiet contributors are visibly underrated, or you genuinely can't tell which parts of a team are struggling. At this point some structure would help, and the question is how much counting you can introduce without buying more distortion than you're solving.

The edge case that forces it. Measures are being used in decisions about pay, promotion or continued employment, or they involve monitoring of some kind. What's permitted, what notice is required, and how measures may be used in employment decisions differ sharply by jurisdiction, and monitoring rules in particular have been changing. Establish what applies where you operate and take local advice before introducing anything.

Five Questions This Reader Asks at 11pm

How do I measure performance objectively? You can't, entirely, and the attempt is where most of the damage happens. What you can do is make judgement more consistent and better evidenced, which addresses the actual complaint. A number isn't objective because it's a number: somebody chose which thing to count, and that choice contains all the subjectivity you were trying to remove, now hidden.

How many metrics should one person have? One or two, plus judgement about the rest. Every additional measure divides attention and adds a distortion, and a person with six is being asked to optimise six things simultaneously, which nobody does. They'll pick the two that are watched most closely, and you've effectively chosen those two by accident.

How do I measure work that isn't countable? By describing the outcome you'd recognise rather than finding something adjacent to count. A great deal of valuable work is judgement, maintenance or prevention, and the countable edges of those jobs are the least important parts. Forcing a number onto them measures the wrong thing precisely.

Should metrics be tied to pay? Be careful, because the strength of the distortion scales with the consequence. A measure that's merely observed produces mild optimisation. The same measure attached to somebody's money produces determined optimisation, including in ways that damage the thing the measure was standing in for. If you connect them, choose measures that are hard to move by any route except doing the job well.

What do I do when a metric gets gamed? First, stop calling it gaming, because in most cases the person is doing exactly what the measure rewards. The useful response is to treat it as information about the design: a measure that can be moved without doing the work was always going to be. Changing it is the answer, and blaming the individual usually isn't.

Three Honest Categories the Approaches Split Into

Count what the systems record. Use whatever your existing tools already produce: tickets, calls, transactions, response times. It's right when the recorded thing genuinely is the job, which happens in a narrow set of roles, and it's cheap because the data exists. It fails almost everywhere else, because the things systems record are the things that were easy to instrument rather than the things that matter. You end up assessing people on the byproduct of your tooling choices, and the resulting behaviour follows the measure rather than the work.

Measure outcomes rather than activity. Assess the result the person was trying to produce rather than what they did. It's right in principle and it's the direction most measurement advice points, because outcomes are what you actually want. It fails on attribution. Outcomes are usually produced by several people plus circumstances, so holding one individual to one is either unfair when things go badly or unearned when they go well, and both errors damage trust in the measure.

Structured judgement with evidence. No metric, but an assessment against stated criteria, evidenced by specific examples. It's right for most roles where the work is judgement, coordination or prevention, and it addresses the fairness complaint more honestly than counting does, because it makes the basis explicit rather than hiding it inside a chosen number. It fails on effort and on comparability. It takes real time, and comparing two structured judgements across managers is harder than comparing two numbers, which is precisely why organisations reach for numbers.

Five Diagnostic Questions You Can Self-Assess Against

For each measure, say what behaviour it will produce. Do it before introducing anything, and be specific rather than optimistic. Every measure has an answer to this question, and the answer is usually visible in advance to anybody who does the job. Asking the people who'll be measured is the cheapest way to find out.

What fraction of the job does this measure cover? Not as a number, but honestly: is the countable part the main part or the edge? Where the measure covers a minority of what somebody does, assessing them on it tells them which minority you value, and they'll act accordingly.

Can this be moved without doing the work? Try to think of three ways. If you can find them easily, so can somebody whose pay depends on it, and that's a design problem rather than a character problem. The measures worth using are the ones that are hard to move except by doing the job properly.

What does this measure say about somebody having a bad month? Test the measure against a realistic bad case: illness, a difficult project, somebody covering for a colleague. A measure that reads a reasonable period of disruption as poor performance will produce exactly one behaviour, which is people hiding the disruption.

Who's advantaged by this measure? Look at who it favours structurally. Measures based on volume favour people with predictable workloads. Measures based on visible output disadvantage those doing preventive or support work. Neither effect is intended and both are reliable, and they compound through pay decisions over years.

Six Kinds of Performance Measure, Reviewed

Output volume

Counting units produced: tickets, calls, transactions, items. It earns its place where the units are genuinely comparable and where volume is the point, which is a narrower set of jobs than it's applied to. It's also easy to collect, unambiguous, and simple to explain, which are real advantages.

Where it falls short is that units are rarely comparable. One difficult case and five simple ones count the same, so the measure quietly rewards selecting easy work, and the people who take the hard cases are penalised for it. It also pushes against quality and against helping colleagues, since neither adds to your count.

If you use it, pair it with something about difficulty or quality, and expect to spend real effort on the definition of a unit, because that definition is where all the behaviour comes from.

There's a second-order effect worth watching in any team where work is allocated rather than chosen. If people can't select their own cases, volume becomes a measure of what they were given, which makes it partly an assessment of whoever does the allocating. That's worth knowing before you use it in a pay decision, because the person being measured has no control over the input.

Quality or error rate

Counting defects, complaints, rework or failures. It earns its place as the natural counterweight to volume, and it captures something organisations genuinely care about. It's also the measure most likely to reflect the thing a customer experiences.

It falls short in a way that's worth understanding, because it's the most damaging in this list. A measure of errors produces an incentive to avoid detected errors, which is not the same as avoiding errors. Depending on the setting, that means avoiding difficult work, not reporting mistakes, or resolving things quietly rather than raising them. In safety-relevant environments that's a serious problem, because visible error rates falling can mean reporting has stopped rather than errors have.

Where you measure errors, measure them at the team level rather than the individual level wherever you can, and make reporting an error something that's never personally costly.

The distinction that matters is between an error rate and a reporting rate, and most organisations can't tell them apart from the data alone. One useful check is to look at how errors are being found: a healthy operation catches most of its own before anybody outside notices, and a rising proportion discovered externally while the internal figure falls is the clearest available sign that reporting has stopped rather than performance improved.

Speed and responsiveness

Time to respond, time to resolve, turnaround. It earns its place where waiting is the customer's main complaint, and it's one of the few measures where improvement is felt immediately by somebody outside the organisation.

It falls short by being the easiest measure in this list to satisfy without doing anything. Responding quickly is trivially achievable: an acknowledgement stops the clock. Resolution time is better and still gamed by closing and reopening, by narrowing what counts as resolved, or by pushing work to a colleague. It also directly penalises the person who takes on the complicated problem.

Use it sparingly, define carefully what stops the clock, and never use it as the only measure in a role where difficulty varies.

It also interacts badly with anything requiring thought. Where the clock is visible and running, people answer before they've understood the problem, because a fast wrong answer scores better than a slow right one. In work where the first response shapes everything after it, that's an expensive trade, and it shows up downstream as rework rather than on the dashboard where the improvement was recorded.

Outcomes the person only partly controls

Revenue, retention, project delivery, results. It earns its place because outcomes are what the organisation actually wants, and measuring activity instead is how you end up with busy people and no results.

It falls short on attribution, and the failure runs both ways. Somebody whose numbers are good may have inherited a strong position, and somebody whose numbers are poor may have been handed a difficult one, and the measure can't tell. Over a single period the noise frequently exceeds the signal, which means you're rewarding and penalising variation rather than performance. People also respond by managing the conditions rather than the work: negotiating targets, choosing favourable assignments, and avoiding anything risky.

If you use outcome measures, look at them over longer periods and alongside a judgement about the hand somebody was dealt.

The cleanest version compares somebody against the position they inherited rather than against a common target. That's harder to compute and considerably fairer, because it assesses the difference the person made rather than the territory they were given. It also removes most of the incentive to negotiate for the easy assignment, which is one of the more corrosive behaviours outcome measurement produces.

Peer and stakeholder assessment

Asking colleagues or internal customers how somebody is doing. It earns its place by capturing what nobody else can see, particularly collaboration, reliability and whether people want to work with somebody again, which is genuinely important and invisible to most metrics.

It falls short when it becomes an input to pay. Once assessments affect money, they stop being honest: people trade favourable ratings, avoid criticising anybody they depend on, and become careful in ways that empty the exercise. It also measures social visibility to a degree that's hard to separate out, so people who work closely with many colleagues do better than equally effective people who don't.

It's most useful as development input and least useful as a decision input, which is the reverse of how it's usually deployed.

One design choice makes a substantial difference, which is who selects the respondents. Where the person being assessed chooses freely, the sample is friendly and the exercise is pleasant and uninformative. Where the manager chooses alone, it can look targeted. Agreeing the list together, with a requirement to include people the work genuinely depends on rather than only those it's enjoyable to work with, produces something closer to useful.

Manager judgement with no metric

An assessment based on the manager's observation of the work, evidenced with examples. It earns its place because it can weigh the things no measure captures: difficulty, context, what somebody was actually asked to do, and the value of work that leaves no trace. For most complex roles it's the only approach that describes the job.

It falls short on consistency and on defensibility. Two managers assess the same work differently, judgement is affected by visibility and by how well somebody presents their own case, and when a decision is questioned an opinion is a weaker position than a record. Those are real problems and they're the reason organisations reach for numbers.

The version that works is structured: stated criteria, specific evidence, and a calibration step. That keeps the ability to weigh context while addressing most of the consistency objection. The evidence requirement is the part that does most of the work, and it is also the part most often dropped for time. An assessment supported by three specific instances is a different document from one supported by a manager's general recollection, both in how fairly it lands and in whether it survives being questioned.

The Decision Table

Situation Scale Setup Primary Pain Recommended Starting Point
Judgement trusted, decisions explainable Any Any None Leave it alone
Two managers assess similar work differently Any Any Consistency, not measurement Stated criteria and calibration first
Units genuinely comparable, volume is the job Any Transactional work Little, if the definition is tight Volume, paired with a quality measure
Difficulty varies enormously between cases Any Any Volume will reward easy work Do not count units. Judge the work
Work is prevention or maintenance Any Any Nothing countable represents the value Describe the outcome, use judgement
Errors are being measured individually Any Safety-relevant Reporting will fall before errors do Move to team level, decouple from consequence
Outcomes depend on many people Any Any Attribution is unfair in both directions Longer periods, plus context judgement
Peer input feeds pay Any Any Honesty disappears from the input Use for development, not decisions
Measures feed employment decisions Any Any Local requirements apply Establish them, take advice first

The fourth row is the most common mistake in this area and the easiest to avoid. Where case difficulty varies widely, any count of cases is a measure of which work somebody chose, and the people who take the difficult work will always look worse than the people who don't.

What Each Measure Will Distort

Every measure produces a predictable behaviour. The useful discipline is naming it in advance and deciding whether you can live with it.

The measure The behaviour it produces What to pair it with
Volume of output Selecting easy work, avoiding complex cases A difficulty weighting, or judgement about case mix
Error or defect rate Avoiding detected errors, which includes not reporting them Team-level measurement, and a safe reporting route
Speed of response Acknowledging fast, resolving slowly A resolution measure, and a clear definition of done
Time to resolve Narrowing what counts as resolved, passing work on Quality or repeat-contact rate
Revenue or results Negotiating easier targets, avoiding risk Judgement about the starting position
Utilisation or billable hours Recording generously, avoiding unbillable but useful work Honest categories for non-billable time
Peer ratings Reciprocal generosity, reluctance to criticise anybody Keep it away from pay decisions entirely
Project delivery dates Padding estimates, declaring done early Something about what was actually delivered

The second row is the one with consequences beyond performance management. In any setting where errors matter, individual error measurement reliably suppresses reporting before it suppresses errors, and an organisation that can no longer see its mistakes is in a worse position than one with a visible error rate. That effect is strong enough that it's worth treating individual error metrics as a design decision requiring specific justification.

The sixth row catches out professional services organisations repeatedly. Where utilisation is measured and non-billable time has one undifferentiated category, everything that isn't chargeable looks like waste, including recruitment, mentoring, internal improvement and helping a colleague. People notice what's counted and stop doing the rest.

Measuring Work Where the Output Is Judgement

Some roles produce almost nothing countable, and the countable fragments are the least important part. A senior specialist whose main contribution is making good decisions under uncertainty, an operations person whose success is that certain things don't happen, a manager whose value is in the problems that never escalate.

The first honest move is to accept that the measure will be a description rather than a number. You can state what a good year looks like in terms specific enough that two people would agree whether it happened: the decisions that were made well, the situations that were handled, the problems that were prevented or caught early. That's assessable without being countable, and the specificity is what makes it fair.

The second is to assess the decisions directly. For judgement-heavy work, the useful evidence is a small number of actual decisions examined properly: what the person knew at the time, what they chose, and what happened. That's a richer assessment than any metric and it has a second benefit, which is that it's also a development conversation, since reviewing decisions is how judgement improves.

The third is to ask the people who see the work. Prevention is invisible by definition, so the manager frequently isn't the best observer. Asking is also the only route that surfaces work somebody has been doing quietly for years without anybody upstream realising it was them. The colleagues who would have been affected if something had gone wrong often have a much clearer view, and asking them is more reliable than inferring from a dashboard.

What to avoid is the tempting fourth option: finding the one countable thing at the edge of the job and building the assessment around it. It is tempting precisely because it produces something that looks rigorous and takes an afternoon, whereas describing the outcome properly takes a conversation and produces something that looks soft. It produces a defensible-looking measure, and it tells somebody whose value is judgement that you're counting their meetings. The measure then does what measures do, and within a year you have somebody optimising the edge of their job.

One practical note for anybody being measured this way. If your work resists counting and you're being assessed on something narrow, say so at the point the measure is set rather than at the review. A written note explaining what the measure doesn't capture is the most useful document you can produce, and it changes the conversation from a dispute about your performance into a discussion about the measure. Written at the start it reads as diligence. Produced at the review, after a poor number, the same note reads as an excuse, which is why the timing matters more than the wording.

What to Put in Writing

Measures are usually introduced with a stated purpose and no record of their expected side effects, which is why the same surprises recur.

Artefact Who owns it When it is written What it prevents
What behaviour each measure is expected to produce Whoever introduces it Before it goes live Discovering the distortion a quarter later
What the measure does not capture The manager, with the person When the measure is set A narrow measure becoming the whole job
What counts as a unit, precisely Operations Before counting starts Every argument about the number
Who is structurally advantaged by this Whoever designs it At design Patterns that compound through pay for years
Whether the measure feeds a decision, and which HR Before it affects anybody Strong optimisation attached to a weak proxy
Local requirements on monitoring and use HR, with local advice Before introduction Collecting or using something you should not

The first row costs ten minutes and prevents most of what goes wrong here. Writing down the expected behaviour before launch also creates the thing you'll need later: a record of what you predicted, against which the actual behaviour can be compared honestly.

Questions to Ask Before You Commit

On distortion. What behaviour will this produce? A bad answer is better performance.

On coverage. What proportion of the job does this represent? A bad answer treats the countable part as the job.

On gaming. Can you name three ways to move this without doing the work? A bad answer is that people wouldn't.

On bad months. What does this say about somebody who had a difficult quarter? A bad answer hasn't tested it.

On advantage. Who does this structurally favour? A bad answer is nobody.

On consequence. Does this feed pay, and is the measure strong enough to carry that? A bad answer is that it's one input among many.

What Getting This Wrong Costs

The first cost is the behaviour you bought without meaning to. A measure introduced for visibility becomes, within a quarter, a description of what the organisation values, and people respond rationally. The support team gets faster and resolves less. The difficult cases wait longer because nobody wants them. None of this is anybody behaving badly, and all of it is expensive in ways that don't appear next to the measure that caused them.

The second cost is what happens to the work nobody counts. Every organisation runs on a quantity of uncounted effort: helping a colleague, noticing something outside your remit, fixing a problem before it becomes visible, mentoring somebody junior. Introduce measures that don't capture any of that and it erodes, slowly, and the erosion shows up much later as a reliability or culture problem that nobody traces back to a measurement decision.

The third cost lands on specific people, consistently. Measures based on visible output systematically disadvantage those doing preventive, supportive or coordinating work, and those roles aren't randomly distributed across a workforce. Where pay and promotion follow the measures, the effect compounds over years, and it looks like a performance difference rather than a measurement artefact, which is precisely what makes it hard to correct.

So before you introduce anything, work out which of three problems you have. A consistency problem means managers assess similar work differently, and the answer is stated criteria and calibration rather than counting. A visibility problem means you genuinely can't see what's happening, and a small number of well-chosen measures will help. A defensibility problem means decisions are being questioned and judgement isn't holding up, and the answer is evidence rather than numbers, which are different things.

When You Are Ready to Go Further

None of this needs a system. It needs each proposed measure tested by writing down the behaviour it'll produce, a written note of what it doesn't capture, and a deliberate decision about whether it should touch pay.

The step beyond introducing measures well is checking them afterwards, which almost nobody does. Six months in, compare what actually happened to what you predicted, and ask the people being measured what they've started and stopped doing. That conversation produces better information about your measurement design than any amount of analysis, and it's the only way to catch a distortion before it's been compounding for years.

HROpsLab publishes independent comparison work across HR tooling, applicant tracking and payroll. We sell nothing, we take no vendor money, and we publish no paid placements. If the next step is looking at what your current tooling makes easy to measure, our comparison work is one place to start.


Frequently Asked Questions

What are performance metrics?

They're the countable indicators used to assess how somebody is performing: output volume, error rates, response times, revenue, utilisation and similar. The important thing to understand about them is that none is neutral. Somebody chose which aspect of a complex job to count, and that choice carries all the subjectivity that measurement was supposed to remove, now hidden inside a number that looks objective. A metric is best understood as a deliberately narrowed view of the work, useful when you've chosen the narrowing on purpose and misleading when the narrowing was determined by what your systems happened to record.

How do you measure performance objectively?

You can't fully, and pursuing it is where most of the harm in this area originates. What's achievable, and what usually addresses the underlying complaint, is making judgement more consistent and better evidenced: stated criteria, specific examples, and a calibration step so different managers apply similar standards. That's not objectivity, but it's fairer than either unstructured opinion or a number chosen because it was available. When somebody asks for objectivity they're usually asking for decisions that can be explained and don't depend on who your manager is, and that request has a better answer than counting.

How many metrics should one person have?

One or two, with judgement covering the rest of the job. Each additional measure divides attention and adds its own distortion, and somebody given six measures won't optimise six things, they'll optimise the two that are watched most closely. That means you've selected those two by accident rather than on purpose, which is worse than choosing them deliberately. The discipline worth applying is that a measure has to earn its place by informing a decision somebody actually makes, and most metric sets contain several that inform nothing.

How do you measure work that isn't countable?

By describing the outcome rather than finding something adjacent to count. You can state what a good period looks like in terms specific enough that two people would agree whether it happened, which gives you assessability without pretending the work is numeric. For judgement-heavy roles, examining a small number of actual decisions works better than any metric: what the person knew, what they chose, what followed. That's a richer assessment and it doubles as development, since reviewing decisions is how judgement improves. What to avoid is measuring the countable edge of an uncountable job.

Should performance metrics be tied to pay?

Be cautious, because the strength of the distortion scales directly with the consequence attached. A measure that's merely observed produces mild optimisation. The same measure determining somebody's income produces determined optimisation, including through routes that damage the thing the measure was standing in for. If you do connect them, pick measures that are hard to move by any means except doing the job well, and avoid anything where you can readily name three ways to improve the number without improving the work. Requirements around how measures may be used in pay decisions differ by jurisdiction, so take local advice.

What do you do when a metric gets gamed?

Stop describing it as gaming first, because in most cases the person is doing precisely what the measure rewards, and framing it as misconduct prevents you from seeing the design problem. Treat it as information: a measure that can be moved without doing the work was always going to be moved that way, and the fault is in the choice of measure. Fix the design, usually by pairing it with something that captures what was lost, or by moving the measurement to team level. Blaming individuals for responding to incentives you created produces resentment and leaves the incentive in place.

How often should you review performance metrics?

Review the measures themselves at least annually, separately from reviewing anybody's performance against them. The specific question is whether each measure is still capturing what it was meant to capture and what behaviour it's actually produced, which you find out by asking the people being measured what they've started and stopped doing. That conversation surfaces distortions far faster than analysis does. Measures also decay: a metric that described the work well when it was introduced can become badly misaligned after a reorganisation or a change in how work arrives, without anybody noticing.

Is manager judgement a legitimate way to measure performance?

Yes, and for most complex roles it's the only approach that describes the actual job, because it can weigh difficulty, context and the value of work that leaves no trace. The objections to it are real: it varies between managers, it favours visible people and confident self-advocates, and it's a weaker position if a decision is challenged. Those are addressable without abandoning judgement. Structured judgement, meaning stated criteria, specific evidence and a calibration step, keeps the ability to assess what matters while answering most of the consistency complaint that pushes organisations towards counting.

Every measure changes the work. The only question is whether you chose the change or discovered it a quarter later.

Share on X Share on LinkedIn

What to do next?

Explore More Articles

Dig deeper into HR Ops strategy, tools, and workflows built for real teams.

Browse the blog →
Join the HROpsLab Community

Connect with People Ops practitioners sharing real workflows, tools, and challenges.

Join now →