Shadow AI at Work: How to Find Out What Staff Are Really Using

Staff are already using consumer AI tools. Six ways to find out what and how much, from OAuth logs to anonymous surveys, and what each one costs you in trust before you learn anything.

Michael Rodriguez Michael Rodriguez 22 min read
Shadow AI at Work: How to Find Out What Staff Are Really Using

TL;DR

  • Core decision: Treat shadow AI discovery as a trust problem first, a security problem second.
  • Do nothing yet: If your sanctioned tool is good and people use it, the hunt creates risk you didn't have.
  • Has to be true: Whichever method you pick, staff have to be able to tell you something useful without fear.
  • How the options split: Cheap and quiet versus accurate and visible. Pick on the relationship you can afford.
  • Decision rule: Start with the least invasive method that can answer the specific question you actually have.
  • What to expect: Better signal than a ban, worse signal than full monitoring, and a clearer picture of where the sanctioned tool is failing.

Priya is the Head of People Operations at a 320-person logistics firm. Last Tuesday she watched a senior account manager paste a draft of a client escalation into a free chatbot to "tighten the wording". The chatbot is consumer, the input is named, the client is in a regulated sector. Priya didn't see this on any dashboard. She heard about it at lunch.

The instinct is to find every instance and shut it down. That instinct will cost you the information you need. The real issue isn't whether staff are using consumer AI. It's why your sanctioned alternative isn't winning the moment it should win on its own.

When You Genuinely Do Not Need to Act on This Yet

Stage one is "the current setup is genuinely fine". You have a sanctioned tool, adoption is real, people use it for the work you expected, and the questions you hear in retros are feature requests, not workarounds. Picture a claims team where the approved assistant is wired into the case system, so the fastest route to a summary is the sanctioned one. Nobody has to choose. The tell is the shape of the complaints: a better export, a longer input limit, a shortcut on a familiar screen. People who have quietly moved elsewhere don't file feature requests. They go silent, and silence reads like satisfaction. If that's you, don't start a discovery project. You will manufacture problems to justify the budget.

Stage two is friction. Staff use the sanctioned tool, then leave for a consumer one when it can't do the job. You hear about it in one-to-ones. Someone in marketing mentions the approved assistant won't take a scanned PDF, so they paste the text somewhere that will. A recruiter mentions the daily cap runs out mid-afternoon. The signal is informal, and it arrives cheerfully, because nobody thinks they're confessing to anything. What you have is a product gap with a stable workaround bolted on. The stability is the worrying part. Action here's product work on the sanctioned tool, not monitoring. Fix the export. Raise the cap.

Stage three is real risk. You handle personal data, contractual data, or regulated material, and the friction in stage two has tipped into routine bypass. The difference is repetition. One person pasting one draft is an incident. A team that has built the paste into its weekly rhythm has a process, and processes are what auditors ask about. Picture an operations lead working out that the monthly client report has been drafted the same way for months, in the same consumer tool, by people who stopped noticing. At this point you owe someone a defensible answer about what is leaving the building. Not a perfect one. A defensible one, meaning you can say what you checked and why you thought that was enough.

Stage four is the edge case. You have evidence of a specific incident, or a regulator has asked. Discovery stops being optional. But by then you also have a much narrower question, and a much shorter list of methods that fit it. Nobody wants a general census. They want to know whether one file, or one class of data, reached one destination, and roughly when. You can scope the work to one team and a window you already know, rather than turning the company into a subject of enquiry. Pressure here pushes towards collecting everything. Collect the answer you were asked for.

Free Weekly Briefing Stay ahead of what's changing in HR and people ops.

Join 4,200+ leaders getting practical insights every week — no fluff, just signal.

Join Free →

Five Questions You Ask Yourself at 11pm

Am I sure they're actually doing this? Probably yes. If your sanctioned tool is slow, capped, or missing features people need, someone is filling the gap. The people who fill it first are under the most time pressure, usually your strongest performers in your worst weeks. Test it cheaply: ask two managers whether anyone has mentioned a tool you don't recognise. Why it matters: certainty without evidence turns into rumour-led policy. You brief the board on something you heard at lunch. Somebody asks how you know, and the credibility you burn is gone for the next thing you need signed off.

What is the worst thing that could be in a prompt right now? A draft email is one thing. A spreadsheet of client names, salary data, or a contract clause is another. Work it out by listing the data classes you'd have to report on if they left, then asking which teams touch them under deadline pressure. The overlap is your answer. Why it matters: the answer sets the urgency and the legal posture. Misjudge it one way and you run a monitoring programme over marketing copy. Misjudge it the other and you stay relaxed about a payroll extract nobody framed as an AI question.

If I find out, what do I actually do with the answer? If the plan is "tell them to stop", you don't need discovery. You need a better tool and a short amnesty. The response you intend should pick the method, not the other way round. A plan that ends in a product fix needs only rough shape, so a survey carries it. A plan that ends in a formal record needs evidence that survives being contested, which is slower and far more expensive. Why it matters: discovery without a response is surveillance with a PowerPoint. If you find things and do nothing, you've taught everyone that the rules are decorative.

What does my IT team already have access to? OAuth grant logs and SSO events often exist before anyone asks. The grant log shows which third-party apps staff connected to a company Google or Microsoft account, and it usually sits in the admin console already, at no extra cost. Identity provider reports and email gateway metadata are often there too. Half an hour with whoever owns the identity stack tells you what's already answerable. Why it matters: the cheapest discovery method may already be paid for. Buying a product that reproduces a report you own spends budget and credibility.

What will staff think when they find out I am looking? They will think you already decided they were guilty. Assume they will find out, because someone always tells, usually whoever configured it. The method arrives as a message before your explanation does. Why it matters: the method you pick is also a message about what kind of employer you're. Get this wrong and the cost isn't a complaint. It's that the next person who nearly sends something they shouldn't says nothing, and you learn it from the client.

Three Honest Categories the Approaches Split Into

Ask, and trust the answer. Anonymous surveys, manager one-to-ones, drop-in sessions. Right when the relationship is intact and you want to fix the sanctioned tool. Fails when staff fear the response, or when the worst offenders never speak up. What this category buys you is the "why", which no log will produce. A log tells you an app was connected. Only a person can tell you they connected it because the approved assistant gives up on anything longer than a couple of pages. Picture a people director who books a drop-in session braced for confessions and gets a queue of complaints about an export limit. The failure mode is subtler than people lying to you: staff answer the question they think you asked. Ask "are you using unapproved tools" and you'll hear no, from people who are. Ask "what did you do the last time the approved tool couldn't finish the job" and you'll get a list.

Read the systems you already own. SSO logs, OAuth grants, identity provider reports, email gateway metadata. Right when you need a defensible record without new spend, and when your identity stack is the source of truth. Fails when people use personal accounts and personal devices on personal networks, which is most of the actual problem. The strength here is that the evidence pre-dates the question. Nobody assembled it to build a case, which makes it hard to argue with later. It's also quiet. An administrator can export a grant list without a colleague noticing. Picture an IT lead who pulls that export expecting chatbots and finds a note-taking app holding mailbox read access, granted long ago by somebody who has since left. Genuine finding. Nothing to do with AI. The limit is the corporate identity itself: the moment somebody signs in with a personal address your logs go blank, and they go blank silently, which is worse than going blank loudly.

Watch the network and the device. CASB, DLP, endpoint telemetry, browser extensions, proxy logs. Right when the risk is high and the data is regulated, and when the previous two categories haven't given you enough. Fails on trust, on cost, and on the basic fact that a determined user on a personal phone is invisible to you. This is the only category that can answer a question about a specific file at a specific moment, which is the question an auditor asks. It's also the only one that can stop something in flight rather than describe it afterwards. Weigh that against what it does to the room. Picture the fortnight after an agent lands on every laptop: nobody objects out loud, and informal reporting quietly falls to nothing. Where it reaches personal devices, the exposure changes shape and becomes a question for local advice rather than procurement. Deploy this because you have a question only it can answer. Never deploy it to find out whether you have one.

Five Diagnostic Questions to Self-Assess Against

What is the single worst category of data that could land in a consumer prompt this week, and can you name it without guessing. Answer it from the data map you keep for your privacy notices, not from the AI conversation. Take the two most sensitive classes on it, then ask which teams handle them under deadline pressure. If you can't produce the name in ten minutes, that difficulty is the finding. A named class fits into a policy sentence a tired person can still remember on a Friday.

How many of your staff have a personal AI account they have used at least once this month, and is your best guess one in three, one in two, or higher. Write the guess down before you look at anything, then check it against the cheapest evidence you have, which is usually the OAuth grant list. The number isn't the point. The gap between your instinct and the record is the point, because that gap is the calibration error you'll carry into every judgement that follows.

What does your sanctioned tool refuse to do that a consumer tool does in ten seconds, and how often does that come up. You answer this by using it, not by asking about it. Take a genuine task off a real team, a scanned document or a long transcript, and finish it inside the approved tool. Time yourself. The moment you feel the pull to open another tab is the moment your staff felt it too, and it contains the whole policy problem in miniature.

If you ran a fully anonymous survey tomorrow, would your most candid answers come from the people who have the most to hide, or the least. Sanity-check it against the last time you asked staff something uncomfortable. Did the free-text boxes fill up, or did you get a wall of neutral scores? Anonymity is a promise people judge by your history, not your wording. If your most exposed teams would stay quiet, the survey isn't wrong, it's just not sufficient. Plan the second method now rather than after the results disappoint.

If you turned on full network monitoring on Monday, would you be able to act on what you found by Friday, or would you be sitting on a problem you can't solve. Test it on paper first. Write the worst plausible finding as one sentence, then the decision it forces and the person who'd have to make it. If that person lacks the authority or the appetite, you've found something worth fixing before you spend anything. Capability you can't act on is just a record of things you knew about and left alone.

The Six Ways to Find Out, Reviewed

An anonymous staff survey

What it is: a short, no-name questionnaire asking which tools staff use and what for. Kept to a handful of questions, it needs no procurement.

Why it earns a place: it surfaces the gap between the sanctioned tool and the actual job, which no log can show. A log tells you a domain was reached. A survey tells you the approved assistant can't read a scanned invoice, which is why the finance team stopped opening it.

Where it falls short: undercounts the worst cases, since the people most exposed are often the least likely to admit it. Anonymity is fragile in practice too. Break the results down by department in a team of six and you have signed every response. What you measure is what people are willing to say, filtered through what they think you want to hear.

Single sign-on and OAuth grant logs

What it is: the record of every third-party app staff connected to a company Google or Microsoft account. Each grant carries a scope, which is the interesting part. An app with read access to a whole mailbox is a different problem from one that can only see a calendar.

Why it earns a place: it's usually already available to IT at no extra cost, and it's the cleanest evidence of which sanctioned-adjacent tools have been blessed. An IT lead can export the list in a morning and sort by scope.

Where it falls short: it misses every personal account used on a personal device, which is where the riskiest inputs often go. It records connection rather than content. A grant made once and abandoned looks identical to one in daily use, so the list overstates breadth while saying nothing about depth.

A cloud access security broker or DLP tool

What it is: a product that sits between users and cloud services and inspects traffic, prompts, and uploads against policy. Some of them block as well as watch, which turns detection into prevention.

Why it earns a place: it's the only category that can see a specific input leaving for a consumer endpoint in near real time. When somebody senior needs to know whether a particular file went to a particular place, this is the only method that answers without asking a human to remember weeks later.

Where it falls short: it's expensive, it's visible to staff on day one, and it doesn't catch personal accounts on personal networks. The cost missed at procurement is tuning. Early alerts are mostly noise, and somebody has to triage them daily. Where nobody owns that queue, the tool degrades into an expensive log nobody reads.

Browser and endpoint telemetry

What it is: agent or browser-level reporting on which domains staff visit and which extensions they install.

Why it earns a place: it closes the personal-device gap that SSO logs leave open, on managed hardware at least. A browser extension with permission to read every page is a live exposure route that rarely appears in anyone's AI policy, and this is the only method here that lists them.

Where it falls short: staff see it as surveillance, BYOD is a legal mess, and it tells you what was visited, not what was typed. A visit to a chatbot domain proves nothing about the prompt. Picture the IT manager who takes his ranked list of heaviest users to HR, where the name at the top is the person trialling a competitor product for a procurement review.

Expense and card reports for personal subscriptions

What it is: a review of corporate card statements and reimbursement claims for paid AI products.

Why it earns a place: paid personal subscriptions are a real exposure vector, and the data already lives in finance. No new tool and no IT time. There's a second signal buried in it: an expensed subscription is a declared one, so the person wasn't hiding, which tells you how safe your staff feel being visible.

Where it falls short: free tiers are invisible here, and staff who expensed a tool aren't necessarily the heaviest users. The free tier is also the one that matters most, since consumer free tiers have historically trained on user input by default while business tiers haven't. So this method reliably finds the safer half of your population and reports back that things look calm.

Simply asking managers in one-to-ones

What it is: a deliberate question added to the existing manager cadence, asked without naming individuals.

Why it earns a place: it uses a channel staff already trust, and it surfaces the workarounds the sanctioned tool is forcing. It costs nothing and it starts this week. It puts the question where the context sits: the manager knows which deadline caused the shortcut and whether it will recur.

Where it falls short: it's anecdotal, managers sanitise, and the worst cases are exactly the ones nobody mentions. Look at the incentive. A manager reporting shadow AI on their team is reporting a gap in their own supervision, upwards, in writing. Picture the team lead who has been drafting her own performance notes in a consumer tool for months. She is not raising it. What reaches you has been filtered twice, and neither filter is malicious.

The Decision Table

Situation Scale Setup Primary Pain Recommended Starting Point
Suspected low usage, no incidents Under 200 staff One sanctioned tool, light adoption Sanctioned tool is probably the problem Anonymous survey plus manager one-to-ones
Regulated data, no current incident 200 to 1,000 staff Mixed identity, hybrid work Defensibility gap if a regulator asks SSO and OAuth grant review, then survey
Known bypass on a specific team Any size Department-level workflow mismatch Repeat exposure on a known data class Manager one-to-ones on that team, then a targeted sanctioned tool upgrade
Incident already happened Any size Any setup Need a record, fast OAuth grant log plus targeted endpoint review, scoped to the team in question
Personal device use is common Over 500 staff, mobile or field roles BYOD, limited MDM SSO logs will miss most of it Manager one-to-ones plus a clear written policy on personal accounts
Audit or regulator has asked Any size Any setup Need a defensible answer with evidence CASB or DLP, scoped narrowly, with a parallel communication plan
You cannot tell whether the risk is real Under 200 staff One sanctioned tool, mixed signals Unclear whether to spend or wait Anonymous survey, with a follow-up commitment in writing
Free tiers are the dominant concern Any size, especially junior-heavy teams Personal accounts in active use Consumer free tiers train on input by default Policy on which login to use for which data, plus SSO enforcement where possible

The Cost of Getting This Wrong

The invoice is the easy part. Network monitoring, a CASB, an endpoint agent, an outside consultant: you can budget those, and a finance director will sign them off without much fight. The harder costs don't appear on any purchase order.

There's the cost of the signal you lose. Staff who feel watched stop telling you things. The next workaround, the next near miss, the next clever prompt that nobody mentions at lunch: all of that goes quiet. You gain a dashboard and lose the only early warning system that ever worked. The loss is invisible while it's happening, which is what makes it expensive. Nobody writes to say they've decided not to mention something.

There's the cost of solving the wrong problem. The moment shadow AI is filed as a conduct issue, the sanctioned tool stops being anybody's priority, because the failure has been reassigned to the people using it. So the product gap survives untouched. The workaround migrates somewhere your chosen method can't reach, usually a personal phone. You've spent the budget confirming a picture that went out of date while you were buying it.

There's the cost of the precedent. A monitoring capability is easy to switch on and awkward to switch off, because somebody then has to argue in writing that the risk has receded, and few people volunteer for that memo. There's the cost to your managers, who now carry a conversation they didn't ask for and can't answer honestly. And there's the version of events that leaves the building with every leaver, repeated in interview rooms you'll never sit in. Obligations you take on when you start collecting are not obligations you can quietly drop, so take local advice before you turn anything on rather than afterwards.

There's also the cost of a punishment delivered in public. Even a quiet one travels. The lesson staff will draw isn't "don't use consumer AI". It's "don't tell anyone when you do". So the question worth sitting with isn't which method is most accurate. It's which method you can still hold a reasonable conversation in six months from now.

When You Are Ready to Go Further

If you've read this far and the picture in your head is sharper than it was an hour ago, that's the most useful outcome. The next step is matching your situation to a small number of credible options, not a long shortlist.

HROpsLab publishes independent, vendor-neutral comparisons of HR and people-ops software, including the AI-tooling category this article sits inside. We don't sell software, we don't resell, and we don't take referral fees from the vendors we cover. Our editorial team writes the reviews and our research team checks the claims, and the two teams report to different leads. If you want a second opinion on the shortlist you're forming, our comparison work is the place to start.


Frequently Asked Questions

Should we just ban consumer AI tools outright?

A blanket ban is faster to write than to enforce, and it's easy to communicate badly. Most organisations that go this route keep the ban in policy and lose it in practice, which is worse than never having banned it. A ban also turns every future disclosure into a confession, so the first thing it costs you is informal reporting. A clearer path is naming which data classes can't leave, which logins are acceptable, and what the consequence is for a real breach. Rules people can follow while rushing are the ones that survive a rush.

What do we do with the information once we have it?

Start by separating the signal from the offender. If thirty people are using a free chatbot to summarise meeting notes, the answer is a better sanctioned summariser, not thirty disciplinary notes. Sort what comes back into two piles: the pattern pile, which is a product and policy problem, and the specific pile, where a named data class left with a named person. Reserve individual action for the cases where a specific data class left with a specific person, and handle those on their own facts.

Is it legal to monitor what staff are doing with AI tools?

It depends on your jurisdiction, your employment contracts, and whether the device is corporate or personal. The shape of the exposure is that you need a clear policy, a clear notice, and a defensible reason, and the bar rises sharply the moment monitoring reaches a personal device or a personal account. Works councils and sector regulators can each change the answer, so an approach that is fine in one country you operate in may not be fine in the next. Take local advice before turning on any monitoring that touches personal devices or personal accounts.

How do we run the survey so people actually answer honestly?

Three things matter. The survey has to be genuinely anonymous, with no email trail and no way to back-trace a response, which in a small team means resisting the urge to break results down by department. The questions have to be about behaviour, not about guilt, so "which tools did you use this month" rather than "did you break the rules". And you have to publish what you'll do with the answers, including the things you won't do. Say that last part before the survey opens, in writing, from whoever would authorise a consequence.

What does our sanctioned alternative have to do to actually win?

It has to do the job a free chatbot does, in the same number of steps, with a privacy story staff can repeat in a sentence. Anything slower, anything more clicks, anything with a quota that runs out on a Tuesday afternoon, and you'll keep losing the moment to the consumer default. The comparison your staff are running isn't between your tool and your policy. It's between your tool and the one already open in the next tab, when they're late for something. The product question is upstream of the policy question.

Do free tiers really differ from paid tiers in terms of what happens to the input?

Consumer chatbot free tiers have historically trained on user input by default, while business tiers haven't. That means the same person, doing the same task, creates different exposure depending on which login they used. So the instruction worth giving staff isn't "don't use AI". It's "use the work login for work", a rule someone can follow while rushing, the only test a rule has to pass. It's one of the few technical facts in this space that has not moved in months, and it's the single best argument for a clear login policy.

How long does a discovery project usually take?

A survey can run in a week. An OAuth grant review is a half-day with the right access. CASB or DLP rollouts run in months and rarely finish on the date the vendor promised, because the real work isn't installation, it's tuning the alerts down to a volume a human will read. Budget for the person who triages that queue, not just the licence. The honest answer is that the time cost scales with how invasive the method is, and the trust cost scales with it too.

What is the single most common mistake?

Starting with the most accurate method, because accuracy is what the board will ask about. Accuracy is the wrong first question. The first question is what you'll do with what you find, and whether staff will still talk to you after you find it. The second mistake is running the exercise once and treating the result as a fixed picture, when what you measured shifts every time a tool ships a feature. Pick something cheap enough to repeat. A rough answer you refresh beats a precise one ageing in a slide deck.

Trust the signal. Keep the relationship. Pick the cheapest method that can answer the question you actually have.

Share on X Share on LinkedIn

What to do next?

Explore More Articles

Dig deeper into HR Ops strategy, tools, and workflows built for real teams.

Browse the blog →
Join the HROpsLab Community

Connect with People Ops practitioners sharing real workflows, tools, and challenges.

Join now →