TL;DR
- The core decision: is the metric something staff would agree to if you asked them in writing, and can they challenge the result.
- When doing nothing is right: when the question is productivity theatre for an exec who already knows the answer they want.
- What has to be true: aggregate, disclosed, challengeable. Take any of those away and the tool decays.
- How the options split: by what is being measured, at what grain, and whether the worker can see and contest it.
- Decision rule: if you can't explain the metric to a person whose pay it touches, in plain words, you don't have a metric.
- What to expect from getting it right: friction, slower rollouts, and the quiet relief of staff who stop rationing their effort to dodge the dashboard.
The Tuesday Morning Memo
Priya, head of people operations at a 600-person fintech, opens her inbox at 8:47. Her COO wants a "productivity heatmap" of the hybrid engineering team by Friday, with names. The request is reasonable on its face. Engineering has missed two releases. The board wants answers. Someone has to produce them.
She closes the laptop for a moment, then opens a vendor tab. Then another, and another after that. The sales decks all say the same thing: visibility, focus, accountability. None of them ask the question she is actually asking, which is whether the heatmap will tell her anything true, or whether it will tell her what the algorithm is paid to tell her, and whether the team will still trust her in six months.
But the deeper issue isn't which tool to buy. The deeper issue is that the request, as written, is a request for surveillance wearing the uniform of analytics, and the difference between those two things is the question this article is built around.
Best tools for AI in the Workplace
When You Do Not Need to Act
Before the framework, an honest map of when this topic is a real problem and when it's a passing pressure that will resolve on its own.
Stage one, your current setup is fine. You have clear goals per role and quarterly reviews that go both ways. Your managers can read a team. Picture a ninety-person design agency where the studio lead can tell you, without opening anything, what each person is working on this fortnight. That is the test. Pass it and a monitoring tool is only a worse copy of what you hold. What the COO usually needs is the last three months of one-on-ones, read back with the delivery dates alongside. Many "we need visibility" requests die on contact with a manager who has been paying attention.
Stage two, friction you can feel. You have a hybrid mix, onboarding is uneven, and managers ask for a sense of whether work is happening. A people partner at a growing insurer notices two new starters in Manchester, six weeks in, on no standing invitation. That came off a calendar, not a surveillance product. This is where lightweight signals earn their place: meeting load, milestone slippage, engagement data staff know they are giving you. Two habits keep it honest. Write down, before you look, what you'll do if the answer is bad. And hold the unit of analysis at team level, because capability ratchets. Own a product that can resolve to a person and every later argument starts from there.
Stage three, real risk. Productivity claims are landing in front of the board. Individual disputes are rising, or there's a defensible concern that a small number of people are coasting. Aggregate measurement is now warranted. Picture an HR director in Dublin assembling a board pack the week before a funding round. She can report delivery against milestones by team and stand behind it. Report it by person and she has created a record that follows the company into every later dispute involving those names. That's the line. Individual measurement without consent, and without works council consultation where the law provides for it, is not the thing she was asked for. Rules differ sharply between jurisdictions, so take local advice early.
Stage four, the edge case. Regulated work, safety-critical environments, customer-facing trust roles with clear fiduciary duty. A dispensing line, a trading desk, a control room at three in the morning. Here some monitoring is part of the job and everyone knows it. The question is still how, not whether. Scope it to the regulated activity rather than to the person, then write down the retention period and who may read the output. Decide what triggers a human going to look. One limit holds even here: emotion recognition in the workplace sits among the prohibited practices in the EU AI Act, so a "wellbeing sentiment" module bundled into a safety product is not a free upgrade. Take local advice.
Five Questions for 11pm
Will I be able to defend this to the person it measures? If no, the tool is producing evidence you can't use. The defence has to work in a small room with one person in it, not in a steering group that already agrees. Write the sentence out. "We measure this because of that, and here is how you can see it and question it." If the sentence needs a slide to survive, it isn't a defence. Get this wrong and you meet the failure late, at the appeal, where you either withdraw the metric and the decision built on it, or defend a number you privately think is noise.
Can the worker see their own data and challenge it? If no, you've built a verdict without a trial. Challenge isn't a courtesy feature. It's the cheapest error detection you'll ever buy, because the person being measured knows about the broken sensor months before the vendor does. She spent Thursday afternoon on a customer call from her mobile and the agent logged her as idle. Without a route for her to say so, the error doesn't get corrected. It gets averaged, and a year later it's a trend in a deck.
Would I be comfortable if my own manager ran this on me? The cleanest gut-check in this space. Use it. Run it against your worst week rather than your best, the one with the funeral and the two nights of no sleep, because that's the week the tool will eventually be read against somebody. If you catch yourself wanting a carve-out for people at your grade, you've designed a tool for a class of person rather than for a purpose. Staff work that out fast. Usually before the pilot ends, and often from the exemption list itself.
What will people do when they know they're watched? They will optimise for the metric. So design the metric to be worth optimising for, or don't measure. This isn't cynicism about staff. It's how any measured system behaves. Mouse movement becomes mouse movement. A document sits open on a second screen all afternoon. Messages get sent at seven in the morning by someone who wrote them the night before. The dangerous part is what gaming does to your confidence rather than your numbers: activity climbs, the dashboard turns green, and the signal you bought stops existing while the reporting grows more assured.
If this leaks, do we survive it? Press, regulator, or a viral internal Slack post. Imagine the screenshot on the front page and decide from there. And assume the screenshot isn't of the data at all. It's of the settings page, or the retention field set to "indefinite", or the internal note explaining which grades are exempt. That note reads differently outside the room where it was drafted. You're answering to two audiences at once, the people who already work for you and every candidate who searches your name before an interview, and only one will ever tell you what they concluded.
The Three Honest Categories
Aggregate, disclosed, challengeable. The workhorse, and the only category that gets easier to defend the longer it runs. Department-level signals, project milestones, engagement data people knowingly volunteered. It's right when you want a sense of a team and not a case file on a person. The useful shape of the finding looks like this: the support function in Berlin sits in twice the meetings of the one in Lisbon and closes fewer tickets, which turns a question about people into a question about a meeting schedule somebody can actually change on Monday. It fails the moment someone asks to "drill down" without a stated, narrow reason. It fails more quietly when the groups are small, because an average of four people is an individual wearing a hat. Before anything is switched on, set a minimum group size and write down who may approve an exception. If the answer is that anyone senior can ask for whatever they like, you don't have a category. You have a filter.
Activity and output, with consent. Time-stamped project artefacts, ticket close rates, code review throughput. Right when roles are well defined and the metric maps to something the business genuinely values, which is most true on support and sales floors where the unit of work is visible and countable by the people doing it. It fails when it treats presence as output. It fails harder on the senior engineer who closes three tickets a week because each one retires a process that used to generate forty, and on the person whose entire quarter went into an incident nobody logged as a ticket. The fix isn't a better algorithm. It's a named human with the standing to annotate the record in writing, whose annotation carries the same weight as the count. Consent matters here in a specific way too. People will accept being counted when they can watch the count as it happens. They will not accept learning that the count existed only when it appears in a review.
Continuous individual surveillance. Keystrokes, screenshots, webcam pings, sentiment scores on private chat. There's a narrow band where a version of this is defensible: regulated activity, high stakes, scope written down in advance, everybody told. In general office work it's almost never the answer. It fails everywhere else, and it fails loudly, because trust doesn't degrade gracefully. It goes in a single meeting. Then metric gaming takes over and the dashboard becomes an accurate record of how hard people are pretending to work. There's a second problem that has nothing to do with culture. Scoring the emotional state of staff from their messages runs close to emotion recognition in the workplace, which the EU AI Act lists among prohibited practices, and performance monitoring of workers is classified as high risk under the Annex III employment category, with those high-risk obligations applying from 2 December 2027 after the Digital Omnibus moved them from 2 August 2026. Take local advice, in every country where you employ people, before anything reaches a shortlist.
Five Diagnostics
Do your managers already know who is carrying and who is coasting? Don't guess. Ask four managers separately, in writing, to name the two people on their team under the most load this month and say why. If the replies come back quickly, with specifics you can check, you have management, and a dashboard will mostly give them something to point at instead of a conversation. If they're vague, you have a manager problem, and the data won't fix it. Either way, more measurement is the wrong medicine.
Can you name the decision this data will change? Write it as one sentence with a named owner and a date. "We will move three people from platform to payments before the quarter closes, and Anna signs that off." Then check Anna has agreed to sign. Half of these requests dissolve at that step, because the owner turns out to be nobody. If the sentence comes out as "we will feel informed" or "we will have visibility", the data is a prop for a conversation somebody doesn't want to have directly.
Have you talked to your works council or staff reps, where the law gives them a say? This isn't a courtesy. In several European jurisdictions it's a requirement, and getting this wrong early is expensive to unwind. Answering it means listing every country where you employ people and finding out, country by country, whether a body with consultation rights exists and what it must be consulted on. Do that before you shortlist, not after you sign. Then ask each vendor which customers went through consultation and what it changed. The ones who have will tell you.
What is your appeals path for an individual who disputes their data? If there's no path, the data isn't safe to act on. Period. Test it rather than designing it on paper. Take one volunteer's real week, produce the report you'd genuinely produce, and have a colleague play the person who disagrees. Time how long it takes anyone in the room to explain why the number is the number. If nobody can trace it back to an event that happened, you're holding a score rather than evidence, and no appeals process rescues that.
Could you publish the metric in the company all-hands, with names attached, and still look your team in the eye? Draft the slide. Not the idea of a slide, the slide, with three real names on it. Read it aloud to one person whose judgement you trust and watch their face while you do. This is the test that catches most bad ideas before they ship, at the price of one uncomfortable afternoon rather than a procurement cycle and a year of damage you diagnose too late.
Six Kinds of Monitoring, Reviewed Honestly
Aggregate collaboration and meeting analytics
What it is: department-level meeting counts, after-hours load, a map of which teams talk to which. Why it earns a place: it surfaces overload and silos without naming names, and answers a question managers cannot see from inside their own team. The useful finding is structural. One team turns out to be the single point of contact for four others, which is why its delivery slips whenever anything else moves. Where it falls short: it treats calendar minutes as a proxy for work, so it rewards the person who accepts every invitation and marks down the one thinking. It is also blind outside the tools it reads, so a team that plans on a whiteboard looks idle in the data. And nearly every product in this class ships a drill-down, so the aggregate framing lasts as long as the first executive who asks for one name.
Individual activity and keystroke tracking
What it is: per-user logs of active seconds, app switches, idle time and typing cadence. Why it earns a place: a narrow set of cases where there's a documented concern about one named person, or a contractual obligation to record activity, with scope and duration agreed before it starts. Where it falls short: the mechanism is broken at the root, because thinking looks exactly like idleness. The person types fewer keys because the problem is hard, the dashboard reports them as coasting, and the company acts on that reading. The agent cannot see the printed spec on the desk or the hour spent on the phone with a customer. You also hold minute-level records of somebody's working day, indefinitely by default, about a person you may later be in dispute with, and that archive gets read back to you by their representative.
Communication sentiment analysis
What it is: scoring the tone of emails, chat and tickets to flag staff who read as negative. Why it earns a place: in theory, a coarse early warning that prompts a conversation a manager would otherwise have missed. Where it falls short: almost everywhere. It marks down the non-native speaker whose English is efficient rather than warm, and anyone having a bad month for reasons that are none of your business. The scores aren't stable enough to act on. Worse, once staff work out that tone is read, tone stops carrying information. Everything flattens into pleasant noise and the early warning you paid for has gone. There's also a legal shape here that the culture argument misses: emotion recognition in the workplace sits among the prohibited practices in the EU AI Act. Take local advice before anyone runs a pilot.
Camera or screenshot capture
What it is: periodic webcam grabs or screen snapshots taken to confirm presence. Why it earns a place: regulated environments with explicit consent and a narrow scope, where the person recorded would describe the reason the same way you would. Where it falls short: this is the one that ends careers and articles. It treats adults as suspects, and it produces screenshots rather than insight, because an image of a screen says nothing about whether the work was any good. It also captures people who agreed to nothing. A kitchen, a housemate walking past, a child, a medical letter open in another tab. Then you store all of it, which turns a monitoring decision into a retention problem. You become the organisation holding a library of photographs of your employees' homes, and you will hold it long after whoever asked for it has left.
Location and badge data
What it is: office entry logs, Wi-Fi association, geofencing. Why it earns a place: it's genuinely good at the questions it was built for. How many desks do you actually need, and who is in the building when the alarm goes off. Where it falls short: presence isn't performance, and the moment badge data gets pointed at people rather than at buildings it becomes a way to punish the carer who works from home two days a week and the person with a two-hour commute. It's also the easiest data in this article to repurpose. It gets bought for space planning, sits in a system with loose permissions, and eleven months later somebody in a leadership meeting asks for a list of everyone attending fewer than two days a week. Nothing technical stops that request. Only a written rule does.
AI note-taking in meetings
What it is: automated transcription and summarisation of calls, sold as a productivity feature rather than a monitoring one. Why it earns a place: the time savings are real, and so are the accessibility gains for a deaf colleague, or for someone reading it back in their third language. Where it falls short: it records who said what, at what timestamp, in a searchable form, which is monitoring whether or not you called it that. The effect on the meeting is what people miss. Once everyone knows there's a transcript, the half-formed idea stops being said out loud, and that idea is what the meeting was for. Article 50 transparency duties under the EU AI Act applied from 2 August 2026, and telling people they're recorded is the floor, not the ceiling. Set the retention period before rollout, because a year of searchable transcripts is a discovery target.
The Decision Table
| Situation | Scale | Setup | Primary Pain | Recommended Starting Point |
|---|---|---|---|---|
| Exec wants a productivity narrative, no specific decision | Any | Centralised HR | Pressure to look informed | Decline, redirect to existing one-on-ones |
| Hybrid teams feel uneven, onboarding is patchy | 200-1,000 | Hybrid, no works council yet | Manager gut feel | Voluntary engagement survey plus manager training |
| Specific team under delivery pressure, names discussed | 100-400 | Hybrid with works council | Need to diagnose a project, not people | Project-level retros and milestone review, no individual monitoring |
| Regulated work, narrow scope, consent in place | Any | Centralised | Audit and compliance | Scoped individual monitoring with documented purpose |
| Layoff round being justified with data | 200+ | Any | Need defensible evidence | Independent legal review before any tool selection |
| Board wants "AI strategy" optics | Any | Any | Nothing operational | No tool. A written policy on what you will not do. |
| Two or more teams pulling in opposite directions | 300-800 | Hybrid | Cross-team friction | Collaboration analytics, aggregate only |
| Sales or support floor, clear output metrics | 100-500 | Onsite or hybrid | Performance management disputes | Existing CRM and QA data, no new surveillance |
| Meeting notetakers already in use, no policy written | Any | Any | Recording by default, consent unclear | Retention rule and attendee disclosure before any new tooling |
The Cost of Getting This Wrong
The invoice is the cheap part. The real costs arrive six months later, and they arrive as second-order effects that no vendor slide warns you about.
First, your best people leave. The ones who can read a job market are the ones you can't afford to lose, and they're the first to recognise a tool that treats them as a case to be closed. You won't get a survey result saying this. Nobody writes "the monitoring" on an exit form, because it sounds petty and they still want the reference. You will get a quiet calendar, then a resignation email with no reason given, then another, then a pattern your dashboard won't catch because the dashboard measures the people who stayed and played along.
Second, the metric becomes the work. Once staff know the score, the score is what they optimise. Tickets get closed quickly and badly. Meetings fill calendars. Chat tone flattens. The dashboard goes green and the work goes grey, and you've built a system that confidently reports on a workplace that no longer exists.
Third, your managers stop managing, which is the cost that compounds furthest out. A manager who can open a dashboard before a one-on-one will open the dashboard, and the half hour turns into a review of the numbers rather than the conversation that would have surfaced the real problem: the caring responsibility she hasn't mentioned, or the offer she's already had. Judgement is a muscle. Outsource it for a year and you're left with a management layer that can only describe its people in the vocabulary a vendor supplied, which is exactly the layer you need working properly on the day the tool gets switched off.
So the question to bring to the executive who asked for the heatmap isn't "do we have the budget". It's "if we run this for a year and the answer is what we want, will the engineering team still be the engineering team."
When You Are Ready to Go Further
If you've read this far, the question is probably no longer whether to act, but which of the small number of options to act on, and in what order. That's the question HROpsLab spends its time on. We are an independent review publication for HR and people operations teams. We don't sell software and we don't sell consulting. We test, we compare, and we publish what we find, including the tools we wouldn't buy.
If you want a structured read on the monitoring, analytics, and engagement category, our comparison work walks through what each approach actually does, what it costs in trust, and where the line sits. If you've a specific scenario and want a second opinion, our team will look at it with you. Either path is a reasonable next move. So is closing this tab and going to talk to your team first, which is often the right answer.
Frequently Asked Questions
Is AI employee monitoring legal?
It depends on where you operate and what you measure, and the rules differ sharply between jurisdictions. In several European countries, works councils have consultation rights over monitoring systems, and some forms of practice are restricted. Performance monitoring of workers is classified as high risk under the EU AI Act Annex III employment category, with high-risk obligations now applying from 2 December 2027, and Article 50 transparency duties applied from 2 August 2026. Emotion recognition in the workplace is among the AI Act prohibited practices. Across borders there is never one answer, only one per country. Take local advice before you act.
Do we have to tell staff we are monitoring them?
In practice, yes, and ethically you should want to. Disclosure is one of the three things that separates measurement from surveillance, along with the ability to challenge and a stated purpose. Hidden monitoring, even where it's technically permitted, produces a workplace you wouldn't want to defend in public. Disclosure means more than a clause in a handbook nobody opened. Say what is collected, who can read it, how long it's kept and which decisions it feeds, in the words you'd use across a desk. If it's drafted so nobody could object, it's cover.
What does "aggregate" really mean?
It means the unit of analysis is a team, a department, or a project, not a person, and the individual isn't identifiable from the output. A dashboard that can be filtered to a single person with one click isn't aggregate. It's individual surveillance with an extra step. Two tests separate the real thing from the label. Set a minimum group size below which the tool reports nothing, then see what happens when a team of five loses two people mid-quarter. And find out who is allowed to break the aggregation.
Do AI meeting notetakers count as monitoring?
If they record who said what, when, and produce searchable transcripts, they're a form of monitoring, regardless of how they're sold. Treat them with the same consent and retention rules you would apply to any recording system, and be clear with attendees about what is captured and who can read it. What settles it isn't whether a person was in the room. It's whether the output can be searched later for what one named individual said. Set retention deliberately, and make it acceptable for anyone to ask for the notetaker off.
What do we do with what we find?
For aggregate signals, you act on them at the team level. For individual data gathered without a documented, narrow purpose and a challenge path, the honest answer is you don't act on it. Data gathered badly is a liability, not an asset, and the safest move is often to delete it and fix the process. The uncomfortable version arrives when the badly gathered data turns out to be right. Act through the route you'd have used before the tool existed, and build a record that stands without it.
How do I answer an executive who wants individual productivity data?
Ask them to name the decision it will change, the person accountable for that decision, and what they will do if the data is contested. If they can answer all three, you've a real request and a narrow path. If they can't, you've a request for reassurance, and reassurance isn't what monitoring tools sell. Put it as a question rather than a refusal, because most of these asks are honest anxiety. Then offer the alternative in the same breath: a milestone review, or a retro on the project that slipped.
Can monitoring ever be the right call for a general workforce?
In narrow, documented, high-stakes contexts with consent and a challenge path, yes. For a general office workforce trying to improve productivity, the answer is almost always no, and the tools that promise to do it well are the ones that fail most loudly in the second-order effects. The distinction that holds up is purpose. Monitoring tied to a regulated activity, scoped to it and visible to the person doing it, survives contact with staff because they'd describe it the way you would. Monitoring a whole population in hope does not.
What is the single best first step?
Write down the decision the data is meant to inform, the metric that would change that decision, and the words you would use to explain the metric to the person it measures. If you can't fill in all three, you're not ready to buy a tool. You're ready to write a policy. Give that policy a section on what you will not do, and get it signed before anyone books a demo. Constraints are far harder to add once somebody senior owns the rollout and has told the board it is happening.
We help HR leaders make sharper decisions on the tools their teams actually use. No vendor relationship, no pay-to-play, no spin.