TL;DR
- The core decision: which format lets you observe the thing you're actually hiring for, in the time a candidate will reasonably give you.
- When doing nothing is right: when the interview already produces evidence you trust, and adding a stage would cost you candidates without changing a single decision.
- What has to be true: somebody has decided in advance what the task is meant to reveal, and what a weak answer looks like as well as a strong one.
- How the options split: by whether the work is observed live, produced alone, or inferred from work already done.
- Decision rule: simulate the conditions of the job, not the difficulty of the job. A task that's harder than the role tells you nothing useful.
- Outcome to expect: fewer stages, a clearer reason for each, and strong candidates who finish the process rather than withdrawing halfway.
The Task Nobody Wants to Set
A hiring manager has two options on the table and doesn't like either. The take-home task the team used last year is thorough, and the last three strong candidates declined it, one of them politely explaining that they'd already done two that month. The live exercise that replaced it is quick and humane, and everybody performs badly in it, including people who turned out to be excellent once hired. The role has been open long enough that somebody has suggested just skipping assessment altogether.
None of these is an unusual position. The take-home that grew from ninety minutes to a weekend, the whiteboard exercise that measures composure rather than competence, the trial day nobody quite knows how to pay for: these are the ordinary ways assessment goes wrong, and they go wrong in the same direction every time, which is toward more work for everyone and less information for you.
The argument is usually framed as accuracy against candidate experience, as though the two trade cleanly and you simply pick a point on the line. They don't. A badly designed take-home is both inaccurate and unpleasant, and a well-designed live exercise can be neither. The real question isn't how much to ask of a candidate. It's what you need to observe, and whether the format you've chosen can actually show it to you.
Best tools for Recruitment & Hiring
When You Genuinely Do Not Need to Act Yet
Your current setup is genuinely fine. The role is one you've hired for repeatedly, the interviewers know what good looks like, and the conversation itself produces evidence you'd defend. Adding an assessment stage here buys you very little and costs you candidates. If your last several hires have worked out and nobody can name a decision the interview got wrong, leave it.
Friction is starting to show. You've had a hire who interviewed well and couldn't do the work, or a debrief where two people disagreed and neither had anything concrete to point at. That's the signal that the conversation has stopped carrying enough evidence on its own. It doesn't yet mean you need a task. It means somebody should write down what the interview is failing to reveal, because that's the thing an assessment would have to show.
It has become a real cost. You're now hiring into a role where a mistake is expensive to unwind, or hiring at enough volume that interviewer judgement varies more than the candidates do. At this point the absence of a common piece of evidence isn't a philosophical gap. It's why your debriefs run long and your decisions correlate with whoever spoke first.
The edge case that forces it. You're hiring for something nobody in the room can assess by conversation, because the work is genuinely technical and the panel is not. Companies hit this when hiring their first specialist in a discipline, and it's the case where an assessment stops being optional. Without one you're evaluating confidence rather than competence, and confidence is exactly the trait that survives an interview best. The same thing happens in reverse when a discipline you do understand starts hiring at a level above anyone currently in the team, because the questions that worked for the previous level stop discriminating.
There's a version of this that catches people out. A process that worked fine while one manager did all the hiring stops working the moment a second manager joins, not because either is wrong but because they're assessing different things and calling them by the same name. The symptom is two hiring managers who both trust their own judgement and disagree about the same candidate, with nothing concrete between them to argue about. That's an assessment problem wearing the costume of a personality clash, and adding a shared artefact fixes it faster than any conversation about alignment will.
Five Questions This Reader Asks at 11pm
What am I trying to find out that the interview can't tell me? If you can't finish that sentence, you're adding a stage out of anxiety rather than need. The answer matters because it determines the format entirely: a question about how someone thinks under questioning needs a live exercise, and a question about the quality of finished work needs a take-home. Get this wrong and you'll run the right process for the wrong question.
Would I do this task myself, this evening, for this role? It's a blunt test and it's the most reliable one available. If the honest answer is no, strong candidates will reach the same conclusion faster than you did, and the ones who accept will be the ones with the fewest other options. That's a selection effect working directly against you, and it operates silently, because the people it filters out never tell you why they stopped replying.
Am I measuring the work or the conditions? A timed exercise under observation measures performance under observation, which is a real skill and usually not the one you're hiring for. If the job involves thinking carefully alone and then explaining it, and your assessment involves thinking aloud at speed while three people watch, you're measuring something the role doesn't require.
Who's going to mark this, and do they agree on what good looks like? An assessment without a marking standard produces the same problem as an interview without one, at greater cost to everyone. If two markers would score the same submission differently, you've added effort without adding evidence, and you've made the candidate pay for it.
What happens to a candidate who does well and doesn't get hired? They've given you real work. If the answer is a template rejection, that's a defensible choice, but make it deliberately. The people most likely to complete a demanding task are the ones most likely to notice how the process ended.
Three Honest Categories the Approaches Split Into
Observed, where you watch the work happen. Live exercises, pair sessions, structured problem-solving with an interviewer present. It's right when the job genuinely involves working with others in real time, when you need to see reasoning rather than output, and when you want to ask why at the moment a decision is made. It fails when the observation itself changes the performance, which it usually does, and it fails hardest for people who work carefully and slowly. It also compresses badly: an hour of watched work is nothing like a day of real work, so you end up assessing a sprint for a role that's a marathon.
Produced, where the candidate works alone and hands something in. Take-homes, written exercises, small projects. It's right when the output is what matters, when the work is genuinely solitary, and when you want something two markers can compare without either of them being in the room. It fails on cost to the candidate, which is entirely theirs and entirely unpaid, and that cost lands unevenly: people with caring responsibilities or second jobs decline more often, so the format quietly filters your pipeline in a direction you didn't choose. It also fails on verification, since you can't be certain who did the work or with what help.
Inferred, where you assess work that already exists. Portfolio review, past-project walkthroughs, reference-led assessment. It's right when candidates already have a body of work, when their context is close enough to yours to transfer, and when you want to spend the time discussing real decisions rather than invented ones. The conversation tends to be richer than anything a set task produces, because the candidate lived with the consequences of what they chose, and you can ask what broke afterwards. It fails when the work is confidential, which is common and legitimate, and it fails when the candidate's contribution is hard to isolate from their team's. It's also the format most likely to reward people who've had good opportunities rather than people who'd do well with one, so it's a poor fit for anyone changing discipline or returning after a break.
A note on combining them. Most teams that run two formats run two from the same category, usually a live exercise and a technical conversation, which is two observations of the same thing at twice the cost. If you're going to run two, take them from different categories. An artefact plus a conversation about it tells you far more than two conversations, and a portfolio review plus one short observed session covers both what somebody has produced and how they think, without asking anyone for an unpaid evening.
Five Diagnostic Questions You Can Self-Assess Against
What did your last bad hire get wrong? Look at the most recent hire that didn't work out and name the specific gap. Was it capability, or judgement, or something about how they worked with people? Then ask whether any assessment format would have surfaced it. Frequently the answer is no, and the honest conclusion is that you have a reference or probation problem rather than an assessment problem.
How long does your process already take? Count the stages a candidate goes through and the calendar time between the first contact and an offer. If you're already at several conversations spread over weeks, adding a task doesn't add rigour, it adds attrition. Ask what you'd remove to make room, and if the answer is nothing, you've learned something about the appetite for this.
Can you describe a weak submission? Not a strong one. Anyone can describe excellence. Write down what a mediocre answer looks like, and what specifically separates it from a good one. If you can't do that before you set the task, your markers won't be able to do it afterwards, and the scoring will drift toward whoever submitted most confidently.
Does the task resemble a day of the job? Compare the assessment to the actual work in three dimensions: how long the person has, whether they can ask questions, and what tools they'd normally use. Every dimension where your assessment differs from the job is a dimension where the result won't transfer. Some difference is unavoidable. All three being different means you're measuring something else entirely.
Who has declined your process recently, and did anyone ask why? Withdrawal is data and almost nobody collects it. If strong candidates are dropping out at the assessment stage, that's the clearest signal available that the format is wrong for the market you're hiring in. It's also the signal that goes unrecorded most reliably, because a withdrawal doesn't generate a debrief. Nobody schedules a meeting about the person who stopped replying, so the evidence that would fix your process is exactly the evidence your process discards. One line in the record, saying at which stage somebody withdrew and what they said about it, costs nothing and is worth more than most of what you collect about the candidates who stayed.
Six Assessment Formats, Reviewed
The take-home task
The candidate is given a defined piece of work and a deadline, and completes it alone in their own time. It earns its place because it produces a comparable artefact: two candidates given the same brief hand in two things you can put side by side, which is the closest hiring gets to a controlled comparison. It also lets people work the way they normally work, with their own tools and at their own pace, which for most roles is a much better simulation than anything observed.
Where it genuinely falls short is the cost, which is entirely the candidate's and entirely unpaid. That cost isn't evenly distributed: people with less free time decline more often, so the format selects for availability as much as ability. Scope creep makes this worse, because a task described as two hours usually isn't, and candidates who care will overrun. You also can't verify the conditions. Someone may have had help, taken far longer than stated, or used tools you didn't anticipate, and none of that is visible in the submission.
The timed live exercise
A problem worked through in a fixed window with an interviewer present. It earns its place by showing reasoning rather than output. You see how someone approaches an unfamiliar problem, where they get stuck, what they ask, and how they respond when a constraint changes. For roles where the work is genuinely collaborative and fast, that's a closer simulation than anything produced in isolation.
It falls short because observation changes performance, and it changes it unevenly. Some people think fluently out loud and some think carefully in silence, and the format rewards the first regardless of which produces better work. It compresses badly too: a role that involves careful work over days gets assessed as a sprint over an hour. And it's expensive on your side in a way that's easy to overlook, because every candidate consumes an interviewer's full attention rather than a marker's spare time.
The paid trial or short contract
The candidate does real work for a defined period and is paid for it. It earns its place by being the only format that assesses the actual job rather than a proxy for it. You see how someone works with your team, your codebase or your clients, over enough time that first-day nerves stop dominating. Both sides get information, which is why it produces fewer regretted hires than anything else on this list.
The limitation is that most candidates can't accept it. Anyone employed elsewhere would have to resign or take leave to participate, which restricts the format almost entirely to people between roles, and that's a narrow and non-random slice of the market. It's also the format with the most significant obligations attached: how trial work must be paid, what employment status it creates and what records you must keep differ considerably by jurisdiction, and unpaid versions carry real exposure in some places. Take local advice before you run one.
The portfolio and past-work review
A structured conversation about work the candidate has already done, with the artefacts in front of you. It earns its place by costing the candidate almost nothing while producing rich evidence, because the work already exists. Done properly it's not a presentation but an interrogation: what was the constraint, what did you try first, what would you do differently. Those answers reveal judgement in a way a fresh task rarely does.
It falls short when the work is confidential, which is common and entirely reasonable, and refusing to share it says nothing about the candidate. Attribution is the harder problem: in team-produced work, isolating one person's contribution depends on their own account, which is exactly what you're trying to verify. And it favours people whose previous employers gave them visible, self-contained projects, which is a function of opportunity rather than ability.
The structured technical conversation
A deep, specific discussion of the domain, with no artefact produced. It earns its place on efficiency: it costs the candidate one conversation, needs no marking infrastructure, and in the hands of someone who genuinely knows the field it discriminates well. A specialist asking follow-up questions can establish depth in half an hour, because the follow-ups are where shallow knowledge shows.
The limitation is that it depends entirely on the interviewer, and it degrades badly when the interviewer isn't a genuine expert. Without real depth on your side it becomes a vocabulary test, rewarding people who use the right terms confidently. It also produces no comparable artefact, so two candidates assessed by two interviewers can't be compared afterwards, which pushes the decision back into the debrief where the most confident voice tends to win.
Reference-led assessment
Structured conversations with people who've worked with the candidate, treated as an assessment stage rather than a formality. It earns its place because it's the only format that observes sustained behaviour over time, which is what you're actually buying. A specific question to a former manager about how somebody handled a particular difficulty produces information no exercise can.
It falls short on candour and access. References are usually chosen by the candidate, current employers often can't be approached, and what a former employer will say is constrained in ways that differ by jurisdiction. Treat what you're told as one input rather than a verdict, and get local advice on what you may ask and record. It's also slow and hard to schedule, which makes it a poor fit late in a competitive process.
The Decision Table
| Situation | Scale | Setup | Primary Pain | Recommended Starting Point |
|---|---|---|---|---|
| Role you hire repeatedly, good track record | Any | Any | Nothing is failing | No assessment stage, keep the interview |
| Interview evidence contradicts itself | Under two hundred | Any | Debriefs run long, no artefact | One short take-home, tightly scoped |
| Work is genuinely collaborative and fast | Any | In person or hybrid | Need to see reasoning | Timed live exercise, one hour |
| Candidates are experienced with visible work | Any | Any | Process already too long | Portfolio review instead of a new task |
| Panel cannot assess the discipline | Any | Any | Evaluating confidence, not skill | Structured technical conversation with an expert |
| Strong candidates keep declining the task | Any | Any | Assessment is costing you the market | Portfolio review, or pay for the time |
| Mistake is expensive and slow to unwind | Any | Any | Need sustained behaviour, not a snapshot | Reference-led assessment alongside one other format |
Most teams sit in two of these rows at once, and the common error is to run formats from three of them simultaneously. Pick the row that matches your worst current failure, run one format for it, and remove a stage elsewhere to make room.
Designing a Task Worth a Candidate's Evening
A task earns the time it asks for by being about the work rather than about filtering. The difference shows up in a few specific design decisions, and each one has a cheap version that saves you effort and costs you information.
| Design decision | The cheap version | The version worth doing | What it reveals |
|---|---|---|---|
| Where the brief comes from | An invented puzzle | A simplified real problem you've actually solved | Whether they handle your kind of ambiguity |
| How long it should take | An unstated expectation | A stated ceiling, honestly set, and a promise you'll read what fits in it | Whether they can scope and stop |
| What they can ask | Nothing, submit blind | A named person they may ask questions of | How they handle incomplete information |
| What is provided | A blank page | Realistic context and constraints | Judgement, rather than guesswork |
| What is assessed | A general impression | Three named things, decided beforehand | Comparability between candidates |
| What happens after | A yes or no | A conversation about their choices | Reasoning that the artefact alone hides |
The clause about reading only what fits in the stated time is the one that changes behaviour. Candidates overrun because they assume everyone else will, and an explicit commitment to judge on the stated window removes the incentive. It also gives you something to observe: someone who scopes to the constraint and says what they'd do with more time has demonstrated the exact judgement most roles need.
The follow-up conversation matters more than most teams expect. An artefact tells you what someone produced, not why, and the why is where the signal is. Ten minutes asking about a specific choice, and about what they'd change, separates people who understood the problem from people who pattern-matched a solution.
Assessing the Output Without Assessing the Person
Marking drifts. The first submission sets an anchor, later ones get judged against it rather than against the standard, and by the fifth the marker has quietly redefined what good means. This isn't carelessness, it's how comparison works when there's no fixed reference, and it's the reason two markers can rank the same set of submissions differently.
Write the standard before the first submission arrives. Name the three things you're assessing, and for each one describe what a weak, adequate and strong answer contains. Do this from the brief, not from the first thing somebody hands in. It takes twenty minutes and it's the difference between marking and reacting.
| Decision | Weak practice | Better practice | Why it matters |
|---|---|---|---|
| Who marks | Whoever has time | A named marker per criterion | Consistency across candidates |
| When the standard is written | After submissions arrive | Before the task is sent | Stops the first submission setting the bar |
| What the marker sees | Name, CV and submission | The submission alone where practical | Reduces drift toward the familiar |
| How scores are recorded | A single overall impression | One score per named criterion, with evidence | Makes disagreement specific |
| When markers confer | Before scoring | After independent scoring | Stops the first opinion anchoring the rest |
Separating the work from the person where you reasonably can is worth doing, though be honest about the limits. Submissions often carry identifying context, and for many roles the conversation afterwards is part of the assessment anyway. The point isn't a perfect blind, it's removing the easy cues that make a marker feel confident for the wrong reason.
What to Put in Writing
Assessment decisions get made in a meeting and remembered by whoever attended. The next hiring round inherits the format without the reasoning, so the task grows, the marking loosens, and nobody can say why any of it is the way it is.
| Artefact | Who owns it | When it is written | What it prevents |
|---|---|---|---|
| What the assessment is for | Hiring manager | Before the format is chosen | Adding a stage out of anxiety |
| The brief, with a stated time ceiling | Hiring manager | Before first send | Scope creep between candidates |
| The marking standard, with three criteria | Marker | Before submissions arrive | The first submission setting the bar |
| Scores with written evidence | Each marker | Immediately after marking | Debriefs decided by whoever speaks first |
| Whether the candidate was paid, and how | Talent lead | At the point of asking | Disputes about unpaid work |
| What was said to unsuccessful candidates | Recruiter | At rejection | Inconsistent feedback and avoidable damage |
The fifth row deserves attention beyond record-keeping. Whether trial or task work must be paid, and what obligations that creates, differs considerably by jurisdiction, and the answer isn't obvious from first principles. Write down what you decided and on what advice, before the first candidate asks.
Questions to Ask Before You Commit
On purpose. What can't the interview tell us? Which decision would this change? If nobody can name a decision that would go differently, you're adding a stage for reassurance. A bad answer describes the format rather than the gap.
On the candidate's cost. How long does this genuinely take, not how long do we say it takes? Would we do it? A bad answer is that serious candidates won't mind.
On marking. Who marks, against what standard, written when? Would two markers agree? A bad answer is that they'll know it when they see it.
On fairness. Who does this format make it harder for, and is that difficulty related to the job? Every format excludes somebody. A bad answer is that it's the same for everyone, which is true and not the question.
On verification. What are we assuming about how the work was done, and does that assumption matter for this role? A bad answer is a plan to detect assistance rather than a decision about whether it's relevant.
On the exit. What do we tell candidates who did well and weren't selected? Do they get their time acknowledged? A bad answer is the standard rejection template.
What Getting This Wrong Costs
The visible cost is candidates withdrawing, and it's the one you notice least reliably, because a withdrawal is silent. Nobody schedules a debrief about the person who stopped replying. So the format that's costing you the strongest candidates looks identical to a quiet market, and the usual response is to widen sourcing rather than examine the process, which spends money on a problem you already had.
The second cost is false confidence. An assessment produces a score, and a score feels like evidence whether or not it measures anything. A poorly designed task with a confident-looking rubric will make a panel more certain and no more accurate, and that's worse than having no task at all, because the certainty suppresses the doubt that would otherwise have prompted a better question.
The third is compounding process weight. Assessment stages are added after a bad hire and almost never removed after a good run, so processes ratchet in one direction. Each addition was justified at the time. The accumulated result is a process that takes weeks, costs candidates real effort, and produces decisions no better than the shorter version did.
So before you add a stage, ask: are you solving an evidence problem, a consistency problem, or a confidence problem? Evidence means you genuinely can't see something. Consistency means you can see it and your interviewers disagree. Confidence means you can see it, you agree, and you're nervous anyway. Only the first two are fixed by an assessment.
When You Are Ready to Go Further
The work above needs no tooling and no budget. It needs somebody to write down what the assessment is for, set a standard before the submissions arrive, and remove a stage when they add one.
The next step, for a reader who wants to know whether their process is unusually heavy for the roles they hire, is comparison against how similar organisations actually run this. That's the part you can't see from inside your own process.
HROpsLab publishes independent comparison work across HR tooling, applicant tracking and payroll. We sell nothing, we take no vendor money, and we publish no paid placements. If the next step is to test your process against the wider market, our comparison work is one place to start.
Frequently Asked Questions
When is a skills assessment worth running at all?
An assessment earns its place when you can name something the interview genuinely can't tell you, and when that thing would change your decision. If your recent hires have worked out and nobody can point to a decision the conversation got wrong, adding a stage costs you candidates and buys you reassurance rather than information. The clearest case for running one is a role nobody on the panel can evaluate by conversation, because without an artefact you're assessing how confidently somebody talks about the work rather than whether they can do it.
How long should a take-home task take?
Short enough that you'd do it yourself, this evening, for this role. That's a better test than any fixed number, because the honest answer varies with seniority and with how competitive the market is for that skill. Whatever you decide, state the ceiling explicitly and commit to judging only what fits inside it. Candidates overrun mainly because they assume competitors will, and a clear promise that you'll read what fits in the stated window removes that pressure while giving you something useful to observe: whether they can scope work and stop.
Should I pay candidates for take-home work?
Paying is defensible and increasingly common, particularly where the task is substantial or resembles real deliverable work. It changes the relationship: you're buying an hour of someone's expertise rather than asking for it. Be aware that the rules here aren't uniform. Whether payment is required, what employment status it may create, and what records you must keep differ considerably by jurisdiction, and unpaid trial work carries genuine exposure in some places. Decide deliberately, take local advice before running anything resembling a trial, and write down what you decided and why.
Take-home or live exercise, which is better?
Neither is better in general, and the choice should follow the work rather than preference. If the job involves thinking carefully alone and producing something finished, a take-home simulates it more honestly. If the job involves reasoning aloud with colleagues under time pressure, a live exercise does. The common error is choosing the format that's convenient for the panel rather than the one that matches the role, then concluding that candidates performed poorly when what actually happened is that you measured a skill the job doesn't require.
How do I stop candidates using AI tools on a take-home?
Start by asking whether it matters for this role. If the job involves using those tools daily, prohibiting them makes your assessment less like the work rather than more. Where it genuinely matters, the practical answer isn't detection, which is unreliable and adversarial. It's designing tasks that depend on context only you have, and following the submission with a conversation about the choices made in it. Someone who produced work they don't understand becomes obvious in about five minutes of specific questioning.
Should every candidate get the same task?
Yes, wherever you can manage it, because comparability is most of the value. Two candidates given different briefs produce two things you can't put side by side, which pushes the decision back onto impressions. The reasonable exception is adjusting the format rather than the substance, for accessibility or for genuinely different circumstances. Keep the criteria identical even when the delivery differs, and record what was varied and why, so that a later reviewer can see the assessment was consistent in the thing that mattered.
Who should mark the submissions?
Somebody who can do the work, marking against a standard written before the first submission arrived. The standard is the part that gets skipped, and skipping it is why marking drifts: the first submission becomes the anchor and everything afterwards is judged against it. Where you have more than one marker, have them score independently before conferring, because a first opinion spoken aloud tends to become the group's opinion. Assign criteria to markers rather than having everyone assess everything, which produces sharper evidence and less duplicated effort.
What do I do when a strong candidate refuses the task?
Treat it as information about your process rather than about them, particularly if it keeps happening. Ask what they'd be willing to do instead, and have an alternative ready: a portfolio conversation, a walkthrough of past work, or a shorter version focused on the one thing you actually need to see. Refusing to accommodate is a defensible position if the assessment is genuinely load-bearing, but be honest about the trade. If your strongest candidates decline most often, the format is selecting for availability rather than ability, and that's working directly against you.
Assess the work you're hiring for, in the time a candidate will actually give you.