TL;DR
Core decision: Choose a scorecard format that matches what your interviewers can actually judge, not what your process wishes they could. Do nothing when: Your team already writes specific behavioural notes and debriefs run on evidence, not averages. Must be true: Every criterion has a written anchor two interviewers would score the same way without conferring. Options split: Numeric scales, anchored behavioural scales, yes-no with evidence, comparative ranking, and single hire/no-hire calls each solve different problems. Decision rule: If interviewers can't describe the behaviour that earns a score, the scale is theatre. Outcome to expect: Debriefs that start with "what did you see" instead of "what did you give them."
The Monday Morning Pipeline Review
A talent lead opens the debrief calendar and sees four scorecards for the same senior engineer candidate. Each one shows four out of five on technical depth, four out of five on communication, four out of five on collaboration. The comments field on three of them says "strong candidate" or "good fit." The fourth is blank. The debrief starts in twenty minutes. The hiring manager has already decided to extend an offer based on the first conversation. The scorecards will be filed. No one will ask why the numbers are identical.
But the problem isn't that interviewers are lazy or untrained. The problem is that the form asked for a judgement on "technical depth" with a five-point scale and no description of what a three looks like versus a four. The interviewer had forty-five minutes with the candidate, a blank box, and a meeting to get to. They picked the number that felt safe and moved on. The real issue isn't compliance. It's that the scorecard launders a gut call into something that looks like evidence.
Best tools for Recruitment & Hiring
When You Genuinely Do Not Need to Act Yet
Your current setup is genuinely fine when interviewers already write down what the candidate said and did, then bring those notes to a debrief where the conversation starts with evidence and ends with a decision. The scorecard exists as a prompt, not a verdict. A talent lead at a forty-person product company describes their process: two interviewers, each with a one-page guide listing three criteria and a blank column for observations. They fill it during the conversation. The debrief is thirty minutes of reading quotes back to each other. No numbers appear. The hire/no-hire call happens in the room. That team doesn't need a new form. They need protection for the time the debrief takes.
Friction appears when the form grows but the conversation doesn't. A hiring manager at a two-hundred-person fintech adds a fourth criterion after a bad hire. Then a fifth after another. The scorecard becomes a checklist. Interviewers tick boxes during the interview instead of listening. The debrief turns into a comparison of totals. The talent lead notices that every candidate scores between seventeen and twenty-two out of twenty-five. The spread is too narrow to decide anything. The form has become a ritual that slows the interview without improving the decision. The fix isn't a better scale. It's cutting criteria until each one matters enough to change the outcome.
Real risk arrives when the scorecard becomes the only record anyone reads. A recruiter at a five-hundred-person retailer pulls the packet for a declined candidate who has filed a complaint. The file contains five completed scorecards. Every criterion is rated four. Every comment field says "met expectations." There's no description of a single answer the candidate gave. The recruiter can't explain why this candidate was declined while another with identical scores was hired. The exposure isn't the scorecard. It's the absence of any evidence the scorecard was meant to capture. The team needs a form that forces evidence into the record before the debrief starts.
The edge case is the team that has already fixed the form but not the culture. A talent lead at a hundred-person SaaS company rolls out anchored behavioural scales with written examples for every level. Interviewers attend a ninety-minute calibration session. The first month produces rich notes and real score variation. The second month the calibration fades. Interviewers revert to fours and fives because the debrief still rewards consensus over disagreement. The hiring manager still decides before the debrief. The scorecard looks better. The decision has not changed. This team doesn't need a new form. They need a debrief structure that makes the scores useful.
Five Questions This Reader Asks at 11pm
How many criteria can I actually score in one interview? Three. Maybe four if the interview is sixty minutes and the interviewer writes nothing else. Each criterion needs a behavioural anchor the interviewer can recognise in real time. Five criteria means five things to watch for, five things to note, five judgements to make while also building rapport and answering the candidate's questions. The interviewer will default to the same number for all of them. Cut until it hurts.
Should interviewers see each other's scores before the debrief? No. The first person to speak in a debrief anchors the conversation. If they have already seen a column of fours, they won't say "I saw something different." They will nod. The debrief exists to surface disagreement. Hide the scores until everyone has stated what they observed. Then reveal. Then discuss.
What do I do about the interviewer who gives everyone top marks? That interviewer isn't the problem. The scale is. A five-point scale with no anchors lets "strong hire" mean "I liked them" and "no hire" mean "I am not sure." Give that interviewer a yes-no question with a mandatory evidence box: "Did the candidate describe a time they debugged a production incident without senior help? Quote the example." They can't inflate a quote. They either have one or they don't.
Can I use the same scorecard for every role? Only if every role requires the same evidence. A sales role needs evidence of pipeline generation. An engineering role needs evidence of system design. A support role needs evidence of de-escalation. A shared "communication" criterion with one anchor serves none of them. Build a library of criterion-anchor pairs. Assemble the scorecard per role. The library is the asset. The assembled form is disposable.
When should the interviewer complete the scorecard? During the interview, not after. Memory degrades fast. The specific phrase the candidate used, the pause before the answer, the question they asked back, all of these vanish within an hour. The scorecard is a capture tool. Treat it like a lab notebook. Write the observation. Score it against the anchor. Move on. If the interviewer waits until the end of the day, the scorecard becomes a justification for a feeling they already have.
Three Honest Categories the Approaches Split Into
Rating scales with defined anchors What it is: A numeric or labelled scale where each point has a written behavioural description. A three on "stakeholder management" reads "describes a situation where they aligned two conflicting stakeholders on a single timeline, naming the compromise each side made." A four adds "and secured written agreement from both." The interviewer matches the observed behaviour to the description. When it works: The criterion is observable, the anchors are specific enough that two interviewers pick the same level independently, and the debrief uses the level as a starting point for "what did you see that made it a three." Where it fails: When the anchor describes an outcome the interviewer can't observe in forty-five minutes. "Builds high-performing teams" isn't observable in an interview. "Describes a hiring mistake and what they changed" is. Teams write the first anchor, realise it can't be used, and revert to gut scoring. Real example: A talent lead at a three-hundred-person logistics company writes anchors for "operational judgement." Level three: "Walks through a past capacity crunch, names the trade-off they made, and states the metric they watched to know it worked." Level four adds "and explains how they communicated the risk to the customer before the miss happened." Interviewers use it. Scores spread. Debriefs reference the trade-off.
Yes-no with mandatory written evidence What it is: A binary question paired with a free-text box that can't be submitted empty. "Did the candidate give a specific example of resolving a conflict with a peer? Paste the example." The interviewer either has the quote or they don't. No middle ground. When it works: The criterion is a threshold behaviour, something the role absolutely requires and the interview is designed to elicit. The interviewer's job is witness, not judge. The debrief reads the quotes and decides whether the pattern meets the bar. Where it fails: When the criterion is genuinely dimensional. "Writes clear code" isn't yes-no. "Writes code with meaningful variable names and extracted functions" can be yes-no if the interview includes a code review exercise. Teams force dimensional criteria into binary boxes and lose signal. Real example: A hiring manager at a seventy-person climate-tech startup uses yes-no for "has shipped a consumer-facing feature end to end." The interview includes a portfolio walkthrough. The interviewer pastes the feature name, the candidate's role, the metric it moved. The debrief has five quotes. Three candidates have them. Two don't. The decision takes ten minutes.
Comparative ranking against a live candidate set What it is: The interviewer ranks the current candidate against the last three they interviewed for the same role on each criterion. "Stronger than candidate A on system design, weaker than candidate B on communication." No absolute scale. When it works: The interviewer has seen enough candidates to hold a mental distribution. The role is high-volume and the same interviewer sees many candidates. The ranking forces discrimination. The debrief aggregates rankings into a relative order. Where it fails: When interviewers see few candidates, or when the candidate pool changes composition mid-process. Ranking candidate five against candidates one through four works. Ranking candidate twenty against candidates fifteen through nineteen when the sourcing channel shifted produces noise. The interviewer can't hold the distribution. Real example: A recruiter at a two-hundred-fifty-person marketplace uses ranking for high-volume sales development roles. Each interviewer sees eight candidates a week. They rank each new candidate against the previous three on "cold call resilience" and "discovery questioning." The debrief sorts by aggregate rank. The top third move forward. The bottom third are declined. The middle gets a second look.
Five Diagnostic Questions You Can Self-Assess Against
When you pull the last ten scorecards for a single role, do the scores vary? Open the ATS. Filter by role. Export the scores. Look at the distribution per criterion per interviewer. If every interviewer gives every candidate a four on "problem solving," the criterion isn't working. It's either unobservable, the anchors are missing, or the interviewer doesn't believe the score matters. Fix the criterion or cut it. Don't retrain the interviewer on a broken form.
Can two interviewers score the same transcript and land on the same number without talking? Take a recorded interview (with candidate consent). Give the transcript to two interviewers who know the anchors. Ask them to score independently. Compare. If they differ by more than one point on a five-point scale, the anchor is ambiguous. Rewrite it. Test again. This is calibration. It takes thirty minutes. Most teams never do it.
Does the debrief start with the scorecard or the notes? Watch the next debrief. Time how long before someone reads a candidate quote. If the first ten minutes are "I gave them a four" and "I gave them a three," the scorecard has replaced the evidence. The scorecard is a index to the evidence. The evidence is the quote. The debrief should sound like "On system design, Sarah noted the candidate chose eventual consistency. I heard them say strong consistency. Let's check the transcript." The scores come after.
When a candidate is declined, can you point to the specific missing evidence? Pick the last three declined candidates. For each, name the criterion they missed and the evidence that was absent. "They didn't give an example of influencing without authority" is a defensible reason. "They scored a three on leadership" isn't. If you can't name the missing evidence, the scorecard didn't capture it. The form failed.
Do interviewers complain about the scorecard after the first week or after the first month? First-week complaints are about familiarity. First-month complaints are about utility. If interviewers still say "this takes too long" or "I don't know what a three means" after four interviews, the form is the problem. Not the interviewer. Not the training. The form. Change the form.
Five Scorecard Designs, Reviewed
The numeric rating scale
A numbered scale, typically one to five or one to seven, applied to each criterion with a label at each end such as "doesn't meet expectations" and "exceeds expectations." It earns its place because every interviewer understands it instantly, it fits in any ATS, and it produces a number the hiring manager can average if they want to. The weakness is specific: without written anchors at every point, the scale measures the interviewer's internal calibration, not the candidate. A four from one interviewer means "would hire." A four from another means "not the best I've seen." The average of those two fours is a fiction. The debrief treats the average as signal. The decision inherits the noise.
The anchored behavioural scale
A scale where every point has a written behavioural description tied to the criterion. Level one: "Cannot describe a relevant example." Level two: "Describes an example but can't name their specific contribution." Level three: "Describes their contribution and the outcome." Level four: "Adds the constraint they worked within and the trade-off they made." Level five: "Adds what they would do differently next time." It earns its place because it turns the interviewer's job into pattern matching rather than judgement. The weakness is specific: writing anchors that are mutually exclusive and collectively exhaustive takes iteration. Most teams write them once, discover gaps during the first five interviews, and never update them. The anchors drift from the role. The interviewer forces the observation into the nearest anchor. The score becomes a proxy for "which anchor was least wrong."
The simple yes and no with written evidence
A binary question for each criterion paired with a mandatory free-text field. "Did the candidate demonstrate experience with Kubernetes in production? Paste the specific project and their role." The interviewer can't submit without text in the box. It earns its place because it eliminates the middle-ground inflation that plagues numeric scales. The evidence is the record. The debrief reads evidence, not numbers. The weakness is specific: it collapses a dimension into a threshold. A candidate who ran a ten-node cluster and a candidate who restarted a pod once both get "yes" if the anchor is "has used Kubernetes." The team must write the anchor at the exact level the role requires. "Has designed a Kubernetes deployment strategy for a multi-region service" is a different threshold. Most teams write the first one. The signal is lost.
The comparative ranking against other candidates
The interviewer ranks the current candidate relative to the previous N candidates on each criterion. "Top of the last four on technical depth. Bottom on communication." No absolute scale. It earns its place because it forces discrimination. Interviewers can't give everyone a four when they must pick a first, second, third, and fourth. The debrief sees a relative order, not a cluster of identical scores. The weakness is specific: it requires the interviewer to hold a stable mental distribution of past candidates. When the interviewer sees one candidate a month, the comparison set is stale. When the sourcing channel shifts, the comparison set is biased. The ranking reflects the pool, not the role. The team hires the best of a weak batch and calls it a win.
The single hire or no-hire call with a written case
The interviewer makes one binary recommendation, hire or no hire, and writes a paragraph explaining the decision referencing specific criteria and evidence. No per-criterion scores. It earns its place because it matches how decisions actually happen. The hiring manager wants a recommendation, not a spreadsheet. The written case forces the interviewer to articulate the reasoning. The debrief compares cases, not numbers. The weakness is specific: it hides disagreement on individual criteria. One interviewer votes hire because of technical depth despite weak communication. Another votes no-hire because of communication despite strong technical depth. Both write compelling cases. The debrief can't see the trade-off because the scores that would reveal it don't exist. The hiring manager picks the case they like. The bias is laundering again.
The Decision Table
| Situation | Scale | Setup | Primary Pain | Recommended Starting Point |
|---|---|---|---|---|
| High-volume role, same interviewer sees many candidates | Comparative ranking | Interviewer ranks each candidate against last three on three criteria | Scores cluster at top, no discrimination | Yes-no with evidence for threshold criteria, ranking for dimensional ones |
| Specialised role, few interviewers, high stakes | Anchored behavioural scale | Three criteria, five-level anchors written from recent successful hires | Interviewers cannot distinguish strong from adequate | Write anchors from actual interview transcripts of recent hires |
| Team already writes good notes, debrief works | No scale | Blank observation column per criterion, no numbers | None, the current process works | Do not add a scale. Protect the debrief time instead |
| Mixed interviewer experience, some new to role | Yes-no with evidence | Five threshold criteria, each with a specific evidence requirement | New interviewers inflate scores, seniors deflate | Mandatory evidence box per criterion, no score field |
| Hiring manager decides before debrief, scorecards filed after | Single hire/no-hire with case | One recommendation, written case referencing three criteria | Scorecards are theatre, decision is made | Require written case before debrief, hide recommendation until all read |
| Multiple interviewers, low calibration, limited time to fix | Numeric with three anchors only | Three criteria, anchors at 1, 3, 5 only, no 2 or 4 | No time for full calibration, need immediate improvement | Define only the floor, the solid middle, and the ceiling |
Writing an Anchor Someone Can Apply
Turning a vague criterion into something two interviewers would score the same way means writing an anchor that describes observable behaviour, not a trait. The table below shows the transformation.
| Criterion | Weak anchor | Usable anchor | Evidence it asks for |
|---|---|---|---|
| Communication | "Communicates clearly" | "Explains a complex technical concept to a non-technical stakeholder using an analogy, checks for understanding, and adjusts if the stakeholder looks lost" | The analogy used, the check-for-understanding moment, the adjustment |
| Problem solving | "Strong problem solver" | "Walks through a past production incident, names the hypothesis they tested first, the data they pulled, and why they ruled it out" | The hypothesis, the data source, the ruling-out reasoning |
| Stakeholder management | "Good with stakeholders" | "Describes a project where two stakeholders wanted opposite outcomes, states the compromise proposed, and names who agreed and who pushed back" | The two outcomes, the compromise, the agreement and pushback |
| Ownership | "Takes ownership" | "Describes a mistake they made in production, the impact, the fix they shipped, and the process change they proposed to prevent recurrence" | The mistake, the impact number, the fix, the process change |
| Technical depth | "Knows the stack" | "Explains the trade-off between two architectural choices for a recent feature, names the constraint that drove the decision, and states the monitoring they added" | The two choices, the constraint, the monitoring |
The usable anchor passes the transcript test. Give the same interview transcript to two interviewers. Both should highlight the same sentences as evidence for level three. Both should agree the candidate didn't reach level four because the monitoring description was missing. If they disagree, the anchor is still ambiguous. Rewrite until they agree. This takes three to four iterations. Most teams stop at one.
The Debrief the Scorecard Is Meant to Serve
The scorecard exists for the debrief. If the debrief doesn't use the scorecard, the scorecard is waste. The debrief structure below assumes a thirty-minute slot, four interviewers, one hiring manager, one facilitator (usually the recruiter).
| Step | Time | Who speaks | What happens |
|---|---|---|---|
| 1. Silent read | 3 min | All | Each person reads every scorecard silently. No discussion. Notes only. |
| 2. Evidence round | 10 min | Interviewers in reverse seniority | Each interviewer reads one quote per criterion. No interpretation. "On system design, the candidate said…" |
| 3. Disagreement surface | 8 min | Facilitator prompts | "Where did we see different things?" Not "who disagrees." The facilitator names a criterion. Interviewers state what they observed. |
| 4. Hiring manager synthesis | 5 min | Hiring manager | Summarises the evidence pattern. States the hire/no-hire lean. Names the risk. |
| 5. Decision | 4 min | All | Explicit call. Hire, no-hire, or seek more evidence. If seek, names the specific criterion and the interviewer who will get it. |
The reverse seniority order in step two matters. The most junior interviewer speaks first. They're least anchored by hierarchy. If the senior engineer speaks first and says "strong on system design," the junior interviewer who saw a gap will hesitate to contradict. Speaking first gives the junior interviewer the floor before the anchor lands. The facilitator must enforce "quote only" in step two. No "I felt they were strong." Only "They said X." The interpretation happens in step three. The hiring manager speaks last in step four because their job is to weigh the evidence, not add more. They haven't interviewed. Their synthesis is the value they add.
What to Put in Writing
| Artefact | Who owns it | When it is written | What it prevents |
|---|---|---|---|
| Criterion-anchor library | Talent lead | Before role opens, updated after each hire | Drift between role needs and interview design |
| Assembled scorecard per role | Recruiter | When role opens, locked before first interview | Interviewers improvising criteria mid-process |
| Completed scorecards | Interviewer | During interview, submitted before debrief | Post-hoc justification, memory decay |
| Debrief notes | Facilitator | During debrief, finalised within one hour | Disagreement about what was said, lost dissent |
| Hire/no-hire rationale | Hiring manager | At decision, before offer | Inability to explain a declined candidate |
| Calibration log | Talent lead | After every calibration session | Anchor drift, interviewer divergence over time |
Written records turn a good decision into a defensible one. The criterion-anchor library is the living document. It grows when a new role needs a new criterion. It shrinks when a criterion proves unobservable. It sharpens when a calibration session reveals ambiguity. The assembled scorecard is a snapshot of that library for one role. Lock it. Don't let interviewers add criteria in the moment. The completed scorecard is the interviewer's witness statement. The debrief notes are the collective sense-making. The hire/no-hire rationale is the decision record. The calibration log is the quality control. Most teams keep the completed scorecards and discard the rest. The discarded ones are the ones that would explain the decision six months later.
Questions to Ask Before You Commit
Anchor specificity Question: Can you show me the written anchor for level three on the most important criterion for this role? Bad answer: "We use a standard five-point scale with labels like 'meets expectations.'" Why it fails: Labels aren't anchors. They describe the interviewer's feeling, not the candidate's behaviour.
Calibration practice Question: When did you last run a calibration exercise where two interviewers scored the same transcript independently? Bad answer: "We did training when we launched." Why it fails: Calibration decays. Without recent practice, interviewers drift to their own internal standards.
Debrief structure Question: Walk me through the first ten minutes of your last debrief for this role. Bad answer: "We go around and share our scores." Why it fails: Sharing scores first anchors the conversation. The evidence never gets spoken.
Evidence requirement Question: What happens if an interviewer submits a scorecard with an empty comments field? Bad answer: "We remind them to fill it out." Why it fails: A reminder after the fact doesn't recover the observation. The form must prevent submission without evidence.
Criterion ownership Question: Who decides which criteria go on the scorecard for a new role? Bad answer: "The hiring manager picks from a list." Why it fails: The hiring manager often picks traits. The interviewer needs observable behaviours. The talent lead must own the translation.
Score visibility Question: Do interviewers see each other's scores before the debrief? Bad answer: "Yes, the ATS shows them." Why it fails: Visible scores create cascade bias. The first score becomes the anchor for everyone else.
Threshold definition Question: For a yes-no criterion, what exactly must the candidate show to get a yes? Bad answer: "Relevant experience." Why it fails: "Relevant" is a judgement, not an observation. The anchor must name the specific behaviour.
Process ownership Question: Who is responsible for updating the anchors when the role changes? Bad answer: "HR updates them annually." Why it fails: Annual updates miss the drift that happens after three hires. The talent lead must own it continuously.
What Getting This Wrong Costs
The first cost is the hire who should not have been hired. The scorecard showed fours across the board. The debrief lasted twelve minutes. The hiring manager extended the offer. Six months later the engineer has not shipped a feature without a senior engineer rewriting the core logic. The team carries the work. The cost isn't the salary. It's the senior engineer's time, the delayed roadmap, the quiet resentment of the team who interviewed the candidate and said nothing because the scorecard gave them no language to say "I saw a gap."
So the second cost is the candidate who should have been hired. The scorecard showed threes because the interviewer didn't know what a four looked like. The debrief averaged the threes. The hiring manager declined. The candidate joined a competitor and shipped the feature your team is still designing. The cost isn't the missed hire. It's the market signal you sent: "We can't tell strong from adequate." The next strong candidate hears it in the interview. They withdraw.
The third cost is the complaint you can't answer. The declined candidate asks for feedback. You have five scorecards with fours and "met expectations." You can't say what they missed. You can't say why the hired candidate was different. The complaint escalates. The legal review finds no evidence in the file. The settlement isn't the cost. The cost is the next hiring manager who learns that the scorecard is a shield, not a tool. They stop trusting the process. They hire on gut. The scorecard becomes theatre for the next cycle.
What would change if the scorecard was the one document you were not afraid to show a candidate?
When You Are Ready to Go Further
HROpsLab publishes independent comparisons of scorecard templates, calibration workflows, and debrief frameworks used by teams that have moved past the cluster-of-fours problem. We don't sell software. We don't run training. We test the artefacts teams actually use and report what holds up under pressure. The comparison library is organised by role type, team size, and interviewer experience level so you can find the starting point that matches your constraint.
If your team has fixed the anchors but the debrief still feels flat, the next piece covers debrief facilitation scripts that surface disagreement without creating conflict. If you're building the criterion-anchor library from scratch, the next piece shows a workshop format that produces usable anchors in ninety minutes with the hiring team. If you're defending the process to leadership, the next piece maps each artefact to the decision it protects.
Every piece is written by practitioners who have sat in the debrief, watched the scores cluster, and rebuilt the form until it worked. No vendor content. No sponsored placements. Just the mechanics that held up.
Frequently Asked Questions
What should an interview scorecard contain?
A scorecard needs the role name, the interviewer name, the date, three to five criteria specific to the role, a scoring mechanism for each criterion (scale, yes-no, or ranking), a mandatory evidence field for each criterion, and a single overall recommendation with a written case. Nothing else. No generic competencies. No culture fit. No "overall impression" field that invites a gut summary. The criteria must come from the criterion-anchor library, not the interviewer's preferences. The evidence field must be required, and the form should not submit without text in it. The overall recommendation is the only synthesis the interviewer writes. Everything else is observation.
How many criteria should I score per interview?
Three. Four at most. Each criterion needs a behavioural anchor the interviewer can recognise in real time. Five criteria means five things to watch for, five things to note, five judgements to make while also building rapport and answering the candidate's questions. The interviewer will default to the same number for all of them. Cut until it hurts. If the role genuinely needs six observable behaviours, split them across two interviews with different interviewers. Each interviewer gets three. The debrief combines the evidence. The scorecard stays usable.
Are numeric scales better than yes and no?
Neither is better. They serve different criteria. Numeric scales with written anchors work for dimensional qualities where the difference between "adequate" and "strong" changes the hiring decision, such as system design depth, architectural judgement, stakeholder nuance. Yes-no with mandatory evidence works for threshold qualities where the role requires a specific experience or not, such as having shipped a consumer feature, has managed a P&L, has debugged a production incident. Forcing a dimensional criterion into yes-no loses signal. Forcing a threshold criterion into a five-point scale creates false precision. Match the scale to the criterion.
When should interviewers complete the scorecard?
During the interview. Memory degrades fast. The specific phrase the candidate used, the pause before the answer, the question they asked back, all of these vanish within an hour. The scorecard is a capture tool. Treat it like a lab notebook. Write the observation. Score it against the anchor. Move on. If the interviewer waits until the end of the day, the scorecard becomes a justification for a feeling they already have. The ATS should lock the scorecard for editing thirty minutes after the interview ends. Late edits are post-hoc rationalisation.
Should interviewers see each other's scores before the debrief?
No. The first person to speak in a debrief anchors the conversation. If they have already seen a column of fours, they won't say "I saw something different." They will nod. The debrief exists to surface disagreement. Hide the scores until everyone has stated what they observed. Then reveal. Then discuss. The ATS should have a "debrief mode" that reveals scores only after the facilitator triggers it. If your ATS can't do this, collect scorecards in a shared folder and open them simultaneously in the debrief.
What do I do about the interviewer who scores everyone highly?
That interviewer isn't the problem. The scale is. A five-point scale with no anchors lets "strong hire" mean "I liked them" and "no hire" mean "I am not sure." Give that interviewer a yes-no question with a mandatory evidence box: "Did the candidate describe a time they debugged a production incident without senior help? Quote the example." They can't inflate a quote. They either have one or they don't. If they still write "yes" with a weak quote, the debrief reads the quote and sees the gap. The evidence does the work the scale failed to do.
How do I score a criterion nobody assessed?
You don't. You leave it blank and note why. If the interview plan assigned "system design" to interviewer A and interviewer A ran out of time, the scorecard shows "not assessed, time overrun on previous criterion." The debrief sees the gap. The hiring manager decides whether to schedule a follow-up or decide on the remaining evidence. Averaging the other scores to fill the gap invents data. The blank is honest. The follow-up is the fix. The process fails when the blank is treated as a zero or a three.
Can candidates request their scorecard?
Retention rules differ by jurisdiction and notes may be disclosable. Take local advice. Assume anything written in the scorecard could be read by the candidate. Write observations you would stand behind in a conversation. "Candidate struggled to explain the trade-off" is defensible. "Candidate seemed junior" isn't. The evidence field protects you. The score field doesn't. If your jurisdiction requires disclosure, the scorecard with written evidence is a stronger position than the scorecard with only numbers. Design for disclosure whether the law requires it or not.
HROpsLab helps hiring teams build processes that hold up under pressure