TL;DR
- The core decision: what sits downstream of the rating, because that's the only thing that justifies compressing a year into a symbol.
- When doing nothing is right: when your ratings feed decisions that genuinely need comparing across managers, and people can explain them.
- What has to be true: you can name the decision the rating serves, and say what would be used instead if it didn't exist.
- How the options split: by how much the scale compresses, and by whether distribution is constrained.
- Decision rule: trace what consumes the rating. If nothing downstream needs comparability, you're paying a real cost for a symbol.
- Outcome to expect: either a defensible scale doing a specific job, or one fewer thing stopping people hearing the conversation.
The Number That Ends the Conversation
A manager has prepared carefully. There's a written assessment with specific examples, a couple of things that went well, one thing that needs to change, and a suggestion about what to work on next. It took an hour to write and it's genuinely useful.
Then the rating gets said out loud. From that moment, the other person is doing arithmetic: what this means for their raise, what it means relative to the colleague who mentioned theirs, whether it's the same as last year. The remaining forty minutes of carefully prepared content lands on somebody who's processing a single symbol, and most of it doesn't register at all.
Both people leave slightly frustrated. The manager feels the preparation was wasted. The person feels they were given a number and a lot of surrounding words. Next cycle, the manager prepares a bit less, because the effort didn't seem to matter, and the year after that the review is mostly the rating with some commentary attached.
Best tools for Performance Management
The real problem isn't that ratings are inherently bad. It's that a rating is a compression, built to serve a decision somewhere downstream, and compression is exactly wrong for a conversation meant to change behaviour. Whether to rate is therefore not a philosophical question about measuring people. It's a practical question about what happens after the rating exists, and most organisations have never traced it.
When You Genuinely Do Not Need to Act Yet
Your current setup is genuinely fine. Ratings feed pay and promotion decisions that need to be comparable across managers, the distribution looks plausible, and managers can explain any individual rating without reaching for impressions. That's what a rating is for, and if it's working, changing it is a project with no problem attached.
Friction is starting to show. Managers are describing the scale as meaningless, or two managers with similar teams are producing visibly different distributions, or somebody has asked what the difference between two adjacent levels actually is. Each of those is specific and each points somewhere. The cheapest check is to ask three managers to define the middle rating and see whether the answers match.
It has become a real cost. A pay decision couldn't be explained, or everybody in a department is rated in the top band, or a promotion was contested on the grounds that the ratings didn't support it. At this point the rating isn't doing the job it exists for, and it's worth working out whether to fix it or drop it.
The edge case that forces it. A rating pattern has been questioned, or ratings feed decisions that are being challenged, or somebody has raised a concern about how ratings distribute across groups. Requirements around fairness, documentation and discrimination in pay and promotion decisions differ sharply by jurisdiction, and rating distributions are one of the places these questions surface. Establish what applies where you operate and take local advice.
Five Questions This Reader Asks at 11pm
Do rating scales actually work? For comparability, yes, which is the job. A scale lets decisions about pay and promotion be made on a common basis across managers who've never met each other's teams, and that's genuinely difficult without one. For development, no, and the mechanism is simple: once a symbol is in the room, it occupies the attention that the rest of the conversation needed.
How many levels should we have? Enough to distinguish what you actually act on differently, and no more. If your pay outcomes have three tiers, a five-point scale is generating two distinctions nobody uses, which produces argument without consequence. Start from the decisions and work backwards, rather than picking a scale and then deciding what the levels mean.
What's forced distribution? Requiring that ratings fit a predetermined shape, so a fixed proportion of people must land in each band regardless of what managers assessed. It solves a real problem, which is ratings inflating until everyone is excellent, by a method most people find objectionable: somebody has to be placed low because the shape requires it rather than because their work did.
Should we abolish ratings? Only if you know what replaces them for the decision they serve. Organisations that remove ratings and change nothing else usually find the decisions move into a quieter, less visible process, which is worse rather than better. The visible symbol was at least contestable.
How do I handle a manager who rates everybody highly? Look at whether that manager is being generous or whether their team is genuinely strong, which is a distinction you can't make from the distribution alone. Then recognise the incentive: rating your own people low costs a manager something immediately, in the conversation and in the relationship, and the benefit is abstract. Without something to counterbalance that, the drift is one-directional.
Three Honest Categories the Approaches Split Into
Rate and constrain the distribution. A scale, with rules about how many people can land in each band. It's right at genuine scale, where pay budgets are fixed and where inflation has already made the ratings meaningless, and it's honest about a trade-off most organisations pretend they aren't making: if the money is finite, somebody is being ranked whether or not you write it down. It fails on what it does to teams and to trust. People discover that their rating depended on their colleagues rather than their work, which is both true and corrosive, and it creates an incentive to avoid working with strong people.
Rate without constraining distribution. A scale, with managers rating as they see fit. It's right in most mid-sized organisations, because it keeps comparability without the visible unfairness of a quota. It fails through inflation. Rating somebody low has an immediate personal cost for the manager and a diffuse organisational benefit, so over a few cycles the distribution drifts upward until the top band is the norm, and the scale has stopped distinguishing anything. Calibration slows this and doesn't stop it.
Don't rate at all. A written assessment with no symbol. It's right where nothing downstream needs comparability, particularly in smaller organisations where whoever makes pay decisions knows the work directly. It's also the only approach that leaves the development conversation intact. It fails at scale and it fails on fairness in an unexpected direction: without a common basis, decisions rest on advocacy and visibility, which disadvantages quiet people and anybody whose manager doesn't push for them. Abolishing the rating doesn't abolish the ranking, it just makes it invisible.
Five Diagnostic Questions You Can Self-Assess Against
Trace what consumes the rating. Follow it downstream. Does it set a pay increase, feed a promotion decision, drive a bonus calculation, or populate a report? If the honest answer is a report nobody acts on, you're distorting every performance conversation in the organisation to produce a number for a spreadsheet.
Ask three managers to define the middle rating. Do it separately. If the definitions differ, the scale isn't comparable, which means it's failing at the only thing it's for. This takes ten minutes and the answers are frequently startling.
Look at the distribution by manager. Compare managers with similar teams. Wide variation is usually about the manager rather than the team, and it tells you the rating currently describes who your manager is as much as what you did.
Look at the distribution across groups. Cut it by tenure, by team, by working pattern, and by anything else you hold. Patterns here are worth understanding rather than explaining away, and they're the kind of thing that attracts scrutiny if a decision is ever challenged. What applies to you differs by jurisdiction, and this is worth advice rather than assumption.
Ask what people remember from their last review. If the answer is the rating and nothing else, that's your evidence for the compression problem. It's not a failure of their attention. It's what a symbol does when it's placed in the middle of a conversation.
Six Approaches to Rating, Reviewed
A numeric scale
Performance expressed as a number, usually on a three to five point range. It earns its place through portability. A number moves easily into pay formulas, into reporting, and into comparisons across departments, and it's compact enough that decision-makers who don't know the individual can still use it.
Where it falls short is in what people do with numbers. A number invites arithmetic that the underlying judgement can't support: averaging across categories, comparing a three to a four as though the gap were meaningful and equal, and tracking movement between years as if it measured something precise. It also tends to acquire cultural meaning that nobody intended, so the middle rating, which is supposed to mean the job is being done well, gets read as a disappointment by almost everybody.
If you use numbers, say explicitly what the middle one means and expect to keep saying it, because the drift towards reading it as failure is constant.
The drift has a cause worth understanding rather than fighting. People don't calibrate against your written definition, they calibrate against what they've heard elsewhere: a previous employer's scale, a colleague's account of their own number, or whatever the middle meant in a school report. You're publishing into a space where the symbol already has a meaning, and yours has to compete with it every year.
Descriptive labels rather than numbers
The same idea with words: exceeds, meets, developing. It earns its place by resisting arithmetic, since you can't average a meets and an exceeds, and by carrying a meaning that's harder to misread than a number on a scale.
It falls short because the words are also compressions and they acquire the same connotations. Developing, in most organisations, means falling short, whatever the definition says, and meets expectations is heard as adequate rather than good. The labels are also harder to use in a pay formula, which is either an advantage or an obstacle depending on what you need.
The honest version pairs each label with a written description of what it looks like in practice, and revisits whether people are using those descriptions or their own assumptions. Wording the labels as statements about the work rather than about the person helps more than the choice of words themselves. A level described as the work met the standard agreed is harder to misread than one described as solid contributor, because the first points at something checkable and the second invites everybody to supply their own scale.
A two-way split
Meeting expectations or not, with no gradations. It earns its place by removing most of the compression damage while retaining the one distinction that genuinely drives different action. In many organisations the real decisions are binary anyway: somebody is fine, or something needs to happen, and the intermediate levels generate argument without changing any outcome.
It falls short where pay genuinely needs differentiating between people who are all performing well. A binary scale can't distribute a variable pay pot, so either the pot becomes flat or the differentiation moves into an unrecorded judgement somewhere else, which is the outcome you were trying to avoid.
It's a strong choice where pay is largely fixed by band and the rating exists to flag exceptions rather than to rank.
The one thing it handles poorly is the very strong performer, who gets the same mark as somebody merely competent and can reasonably ask what the point was. Where that matters, the usual repair is to keep the binary rating for the formal record and handle exceptional contribution through a separate, explicitly discretionary route, which at least stops the scale pretending to a precision it does not have.
Forced distribution across a curve
Ratings must fit a predetermined shape. It earns its place as the only reliable answer to inflation, and it's honest about a constraint that exists regardless: where the money is finite, somebody is being ranked. Making that explicit is at least more transparent than a process that produces the same outcome invisibly.
It falls short on almost everything else. It requires a manager to place somebody low because the shape demands it rather than because their work did, which is indefensible to the individual and known to be indefensible by the manager delivering it. It damages collaboration, since being on a strong team becomes a disadvantage. And where the distribution is applied to small groups, the arithmetic becomes arbitrary in a way everybody can see.
If inflation is the problem, calibration and clearer definitions are worth exhausting first, because this cure is reliably worse than most of the diseases it treats. There is also a softer version that avoids the worst of it: guiding a distribution rather than enforcing one, so managers are told what shape is expected and asked to explain departures rather than being prevented from making them. That keeps the pressure against inflation while leaving room for a team that genuinely is unusually strong.
No rating, with a written summary instead
An assessment in prose, with no symbol. It earns its place by leaving the conversation intact, which is a genuine benefit rather than a soft one: people hear the content, engage with it, and respond to what was said rather than to what it implies about their raise.
It falls short at the point where somebody has to decide between people. Pay pools, promotion slates and calibration discussions all require comparing individuals across managers, and prose doesn't compare. What happens in practice is that decision-makers construct an informal ranking from the summaries, which is a rating with none of the transparency and none of the definitions.
Where this works, it works because something else carries the comparability, usually a separate promotion process with its own criteria.
It also asks considerably more of whoever reads the summaries. Prose takes longer to absorb than a symbol, so a decision-maker comparing forty people will either read carefully, which takes real time, or skim for signal words and construct an informal scale from adjectives. The second is what usually happens, and it produces a rating with no definitions at all, which is the worst available outcome.
Rating the work rather than the person
Assessing specific outputs, projects or objectives rather than assigning an overall level to the individual. It earns its place by being more accurate, since most people are genuinely stronger at some parts of their job than others, and by being easier to hear: a judgement about a piece of work is not a judgement about a person.
It falls short on aggregation. Decisions downstream usually need one answer, so somebody still has to combine the parts, and that combination is either a hidden weighting or an overall rating arrived at by another route. It also produces more assessments to write, which is real work for a manager with a large team.
It's the best fit where the decisions it feeds are about specific capability rather than about overall standing, such as deciding who leads which project.
It has a second benefit that's easy to miss. Rating work rather than people makes it much easier to say something critical, because the sentence is about a piece of work that didn't go well rather than about a person who isn't good enough. Managers who find assessment conversations difficult often discover that this framing alone removes most of what they were dreading.
The Decision Table
| Situation | Scale | Setup | Primary Pain | Recommended Starting Point |
|---|---|---|---|---|
| Ratings feed decisions, distribution plausible | Any | Multi-manager | None | Leave it alone |
| Nothing downstream needs comparability | Under fifty | Single decision-maker | The symbol costs more than it buys | Drop the rating, keep a written summary |
| Three managers define the middle differently | Any | Any | The scale is not comparable | Written definitions, then calibration |
| Everybody lands in the top band | Any | Any | Inflation, and one-directional incentives | Calibration before any forced distribution |
| Wide variation between similar managers | Any | Any | The rating describes the manager | Calibration, with examples not adjectives |
| Pay pool is fixed and differentiation needed | Over one hundred | Variable pay | Ranking is happening either way | Constrain lightly, and say so openly |
| People remember only the number | Any | Any | Compression is eating the conversation | Move the rating out of the conversation |
| Distribution varies across groups | Any | Any | May attract scrutiny | Look properly, take local advice |
| Ratings removed, decisions now invisible | Any | Any | Ranking moved somewhere unaccountable | Give the decision its own explicit process |
The second row is the one most organisations should check and almost never do. Tracing what actually consumes the rating takes an afternoon, and a surprising number of organisations discover it feeds a report that informs nothing, which is an expensive thing to distort every performance conversation for.
What the Rating Is Actually For
A rating only earns its place if something downstream needs it. Working through the decisions it feeds is the fastest way to find out whether yours is doing anything.
| The downstream decision | Does a rating help it? | What would work instead |
|---|---|---|
| Distributing a fixed pay pool | Yes, this is the strongest case | Nothing simpler, if the pool is genuinely constrained |
| Deciding a promotion | Partly, and it is a poor primary input | Promotion criteria assessed directly against evidence |
| Identifying who needs support | No, it is far too coarse | A manager naming the specific gap |
| Identifying flight risk in strong performers | No, the rating tells you the wrong thing | A conversation, which most organisations skip |
| Deciding who leads a project | No, it aggregates away the relevant part | Assessment of the specific capability needed |
| Reporting performance to leadership | Rarely, and this is the weakest case | A written summary of themes, if anybody reads it |
| Supporting a decision if it is challenged | Only if the definitions are real and applied | Contemporaneous evidence of what happened |
The second row deserves attention because it's the most common justification given and it's weaker than it appears. A rating tells you how somebody performed in their current role, and promotion is a question about a different role. Using performance as the primary promotion signal is how organisations end up promoting people out of jobs they were excellent at into ones they're unsuited to, which is a well-known pattern that nobody designs deliberately.
The last row matters for a practical reason. Where ratings are used to support a decision that's later questioned, what carries weight is whether the definitions were real, applied consistently, and evidenced. A rating with no written definition behind it is a manager's opinion expressed as a symbol, and that's a considerably weaker position than a specific account of what happened. Requirements here differ by jurisdiction, so it's worth advice rather than assumption.
Calibration: Useful Meeting, Terrible Substitute
Calibration is the practice of bringing managers together to compare ratings before they're finalised, and it's the standard answer to both inflation and inconsistency. It's genuinely useful and it's asked to do considerably more than it can.
What it does well is expose differences in standard. A manager who discovers that their strong performer would be a middling one in another team learns something that no written definition would have taught them, because the calibration conversation is concrete: somebody describes what a person actually did, and other managers respond with what that would mean in their team. That's the mechanism, and it works.
What it does badly is anything that requires courage in a room. Calibration meetings are social events with a visible hierarchy, and the outcomes reflect that. Managers who advocate confidently get better outcomes for their people than managers who don't, which is a new unfairness layered onto the one being fixed. The most senior person's view tends to settle contested cases. And the people discussed last, when everybody is tired and the shape is nearly agreed, get a materially different quality of consideration.
It also fails when it becomes about the distribution rather than the individuals. Once a group is trying to reach a target shape, the conversation stops being about whether a rating is accurate and becomes about which person can be moved with least resistance, which is usually whoever's manager objects least.
Three things make it work better and none of them is complicated. Discuss examples rather than adjectives, so the conversation is about what somebody did rather than about whether they're strong. Randomise or rotate the order so the same people aren't always discussed when attention has gone. And have somebody explicitly responsible for noticing when a decision is being made on the basis of who argued hardest, which is the failure mode this format produces most reliably.
What to Put in Writing
Rating systems run on shared assumptions that turn out not to be shared, which is why the same arguments recur every cycle.
| Artefact | Who owns it | When it is written | What it prevents |
|---|---|---|---|
| What each level means, in observable terms | HR, with managers | Before the scale is used | Three managers defining the middle differently |
| The decision each rating feeds | HR | Before the cycle | A scale distorting conversations to fill a report |
| The evidence behind each individual rating | The manager | At the time of rating | A symbol standing in for an opinion |
| What was discussed in calibration, and why it changed | Whoever runs it | At the meeting | Adjustments nobody can later account for |
| The distribution by manager and by group | HR | Each cycle | Patterns nobody notices until they are questioned |
| Local requirements on how ratings may be used | HR, with local advice | Before ratings feed pay | A decision that cannot be defended if challenged |
The first row is where most scales fail and it's the cheapest thing on the list. Written definitions in observable terms, tested by asking several managers to apply them to the same described case, will surface disagreement immediately, which is the point.
Questions to Ask Before You Commit
On purpose. What decision consumes this rating? A bad answer is that it goes into the system.
On definitions. Would three managers define the middle level the same way? A bad answer is that it's in the guidance.
On inflation. What's the cost to a manager of rating somebody low? A bad answer is that it's their job.
On distribution. How does this vary by manager and by group? A bad answer is that it hasn't been looked at.
On the conversation. What do people remember from their review? A bad answer assumes they remember the content.
On removal. If you dropped the rating, what would make the decision instead? A bad answer is that managers would work it out.
What Getting This Wrong Costs
The first cost is the conversation you paid for and didn't get. A manager who spends an hour preparing a useful assessment and then delivers a symbol that erases it has wasted most of that preparation, and they learn from it. Over a few cycles the preparation shrinks to match the part that lands, which means the rating gradually becomes the entire content of the review, and the organisation has traded a development conversation for a number.
The second cost is what inflation does to the thing you built the scale for. A scale that drifts until nearly everybody is in the top band no longer distinguishes anybody, so the decisions it was meant to support are made on something else, usually informal advocacy. You're now carrying the full cost of the rating, in damaged conversations and manager time, and receiving none of the comparability benefit.
The third cost shows up in patterns. Rating distributions that vary by group are common, rarely examined, and compound through pay and promotion over years. Whether that pattern reflects something about the work or something about who assesses it is a question worth answering deliberately, because the alternative is discovering it when somebody else asks. What that exposes you to differs by jurisdiction, and it's one of the few places in performance management where the arithmetic is visible enough to be challenged.
So before you change anything, work out which of three things you have. A purpose problem means nothing downstream needs the rating, and the answer is to stop producing it. A definition problem means managers are applying different standards, and the answer is written definitions tested against real cases. A placement problem means the rating is fine and it's sitting in the middle of a conversation it destroys, and the answer is to move it rather than to remove it.
When You Are Ready to Go Further
None of this needs a system change to begin. It needs the downstream decision traced, three managers asked to define the middle level, and the distribution cut by manager and by group.
The step beyond that is deciding where the rating is communicated. Most of the harm attributed to ratings is actually caused by placing them inside the development conversation, and separating those by several weeks costs nothing. An organisation that keeps its scale and moves the delivery has usually solved the problem it thought required abolishing the scale.
HROpsLab publishes independent comparison work across HR tooling, applicant tracking and payroll. We sell nothing, we take no vendor money, and we publish no paid placements. If the next step is looking at what your current tooling assumes about rating scales, our comparison work is one place to start.
Frequently Asked Questions
What is a performance appraisal?
It's the assessment part of performance management: forming a judgement about how somebody has performed over a period, usually expressed as a rating or a written summary, and usually feeding a decision about pay, promotion or development. The word is often used interchangeably with performance review, though they're slightly different things: the review is the process and the conversation, while the appraisal is specifically the evaluative judgement within it. What matters more than the terminology is what the appraisal is for, since an assessment that feeds a real decision needs different properties from one that exists to be filed.
Do performance rating scales work?
For the job they exist to do, which is making judgements comparable across managers who've never seen each other's teams, yes. That comparability is genuinely difficult to achieve any other way and it's what pay and promotion decisions at scale depend on. For development, no, and the reason is mechanical rather than philosophical: a symbol placed in the middle of a conversation occupies all the attention, so the carefully prepared content around it doesn't register. The scale isn't good or bad in itself. It's well suited to one purpose and actively destructive to another, and most organisations use it for both simultaneously.
How many performance rating levels should you have?
However many distinctions you actually act on differently, which is usually fewer than the scale you have. If pay outcomes come in three tiers, a five-point scale generates two distinctions that change nothing, producing argument without consequence and giving managers a harder job for no benefit. Start from the decisions and work backwards rather than choosing a scale and then defining the levels. One practical warning about odd-numbered scales: whatever the definition says, the middle level is read as a disappointment by nearly everybody, so if it's meant to mean the job is being done well, expect to say so repeatedly.
What is forced distribution and should you use it?
Forced distribution requires ratings to fit a predetermined shape, so a set proportion of people must land in each band regardless of what their managers assessed. It solves the real problem of ratings inflating until everybody is excellent, and it's honest about a constraint that exists anyway, since a fixed pay pool means somebody is being ranked whether or not it's written down. The costs are severe: managers place people low because the shape demands it, collaboration suffers because being on a strong team becomes a disadvantage, and in small groups the arithmetic is visibly arbitrary. Exhaust calibration and clearer definitions first.
Should you get rid of performance ratings?
Only if you know what makes the decision instead. Organisations that remove ratings without replacing the comparability usually find the decisions move into a quieter, less visible process, where they rest on advocacy and how well somebody's manager argues for them. That's worse than the thing removed, because at least a rating is explicit and contestable. Abolishing the rating doesn't abolish the ranking. If nothing downstream genuinely needs to compare people across managers, dropping the symbol is a clear improvement, and the test is to trace what currently consumes it.
What is a calibration meeting for?
Bringing managers together to compare ratings before they're final, so that a strong performer in one team means roughly what it means in another. It works when the conversation is about examples rather than adjectives, because hearing what somebody actually did lets other managers say what that would mean in their team, which no written definition achieves. It fails when it becomes about hitting a distribution, at which point the discussion shifts from whether a rating is accurate to which person can be moved with least resistance. Watch the order, too, since people discussed last get noticeably less consideration.
How do you handle a manager who rates everyone highly?
Start by checking whether they're being generous or whether the team is genuinely strong, which the distribution alone can't tell you. Then look at the incentive rather than the individual, because it's one-directional: rating somebody low costs a manager immediately in an uncomfortable conversation and a damaged relationship, while the benefit is abstract and lands on the organisation. Without something to counterbalance that, drift is the expected outcome rather than a failing. Calibration with concrete examples helps most, since it's harder to defend an inflated rating when describing specific work to peers than when entering a number in a form.
Should employees see their rating before the conversation?
There's a genuine trade-off. Sending it in advance means somebody absorbs the number privately rather than in front of the person who assigned it, and arrives able to discuss it rather than react to it, which usually produces a better conversation. The risk is that they read it alone with no context and reach their worst interpretation before you can say anything. If your review and pay discussion are properly separated, this matters less, since the rating conversation is then its own event with a clear purpose. If everything is in one meeting, advance notice mainly moves the shock earlier.
A rating is a compression built for a decision made somewhere else. Find that decision before you defend the scale.