AI HR Tools 24 min read

What Your AI Is Learning From

Three different questions get asked as one, which is why the conversations go nowhere. Six things a feature might be reading, what leaves and what stays, and the prompt box nobody governs.

Sarah Mitchell Sarah Mitchell 24 min read
What Your AI Is Learning From

TL;DR

  • The core decision: what an AI feature reads, and what of yours leaves the building.
  • When doing nothing is right: when you can already answer both, in writing.
  • What has to be true: somebody can say what a feature consumed to produce a given answer.
  • How the options split: by whether the source is yours, the vendor's, or somebody else's entirely.
  • Decision rule: separate what it reads now from what it was built on from what it keeps.
  • Outcome to expect: a clear answer to the question an employee will eventually ask.

Switching It On Reads Everything

The feature is included, somebody suggests enabling it, and the question nobody quite asks out loud is what exactly it's about to read.

It's an awkward question because the answer sounds obvious. It reads your HR data, of course, that's the point. But that answer covers at least three different things, and they have different consequences. Reading a record to answer a question right now is one thing. Having been built on an enormous amount of material before it ever encountered your organisation is another. Your information being used to improve the product for other customers is a third, and it's the one people mean when they get uneasy, and it's the one least often addressed directly.

Those three get discussed as a single topic, which is why the conversations go nowhere. Somebody asks whether the AI is trained on their data, the vendor answers a slightly different question honestly, and everybody leaves with a different impression of what was agreed.

The reframe worth holding: what a feature consumes determines both what it can possibly know and what of yours has gone somewhere. Those are two separate concerns and they need separate answers. The first is about capability, because a feature that can't see your records can't tell you anything specific about your organisation, however good it is. The second is about exposure, and it's the one that ends up in front of an employee asking what happened to their information.

Obligations around personal data differ by jurisdiction, differ by sector, and change. What you may collect, how long you may keep it, what you must tell people, and what happens when information moves between organisations or across borders are all questions with local answers. Nothing here tells you what applies to you. Establish it with local advice before switching anything on.

When You Genuinely Do Not Need to Act Yet

You can already answer both questions. Somebody can say, in writing, what each enabled feature reads and what's retained. That's the state you want and it's rarer than you'd expect.

Free Weekly Briefing Stay ahead of what's changing in HR and people ops.

Join 4,200+ leaders getting practical insights every week — no fluff, just signal.

Join Free →

A feature is enabled and nobody knows what it reads. Common, and worth fixing regardless of whether you change anything else. It's an afternoon's work and it's the foundation for every other decision here.

Somebody has asked and you couldn't answer. An employee, a manager, or somebody doing due diligence. That's the point at which not knowing becomes a live problem rather than a gap.

The edge case that forces it. A team is using something outside the approved tooling, pasting real information into it. That's a different conversation from the one about your own systems, and it's usually more urgent.

Five Questions This Reader Asks at 11pm

Is our data training their model? The question everybody means and the one to ask directly, in those words. The answer may reasonably be no, may be yes with options, or may be that it depends on the feature. What you want is a plain answer plus where it's written down, because a verbal assurance in a sales conversation isn't a position you can rely on later.

What's the difference between training data and what it reads? Training material is what shaped the capability before it ever encountered you. What it reads is the information it consults to answer the specific question in front of it. A feature can read your records without your records having trained anything, and it can also have been built on material you've never seen while touching nothing of yours at the moment of use.

Can it see things the person asking shouldn't see? The question to ask about any feature answering questions across your records. Permissions applied to a person don't automatically apply to a system reading on their behalf, and this is a common gap rather than an exotic one.

What happens to what somebody types into it? Free text is where the unexpected disclosures happen, because people explain their situation. A question about a colleague, a health circumstance, a complaint: all typed in as context, all now somewhere. Worth establishing where that goes before anybody uses it.

Should we worry about our own history? Your accumulated records reflect decisions your organisation made, and a feature drawing on them inherits whatever patterns are in there. That's a real concern and it's a governance question rather than a technical one, so it belongs with the material we publish on AI governance, alongside local advice.

Three Questions People Ask as One

The question as asked What it's really asking Why the answers differ
Does it use our data Which of three things do you mean Reading, training and retention are separate
What is it trained on What shaped it before it met us Frequently nothing to do with your records
What does it read What it consults to answer this question May be live records, may be nothing of yours
Does our data leave Does anything go outside our systems Depends on where the processing happens
Do they keep it What's retained after the answer Retention differs from processing
Does it improve from us Does our material train it for others The question people actually mean
Can it see everything Does it respect the asker's permissions A common and consequential gap
Where does typed text go Free text is data too, and often forgotten Prompts are frequently handled differently

The first three rows are the whole reason these conversations fail. Somebody asks whether the AI uses their data, a vendor answers about training when the questioner meant retention, and both parties finish satisfied and disagreeing. Asking the three separately produces three answers, and the three answers are what you need.

The sixth row is the one people care about and the one phrased most indirectly. Ask it plainly: is anything we put in used to improve this product for anybody else, and where is that written. A vendor with a clear position will tell you immediately.

The seventh row is the practical risk that gets missed while everybody discusses the philosophical one. A feature summarising across records may be reading things the person asking has no right to see, and the output won't say so.

Five Diagnostic Questions You Can Self-Assess Against

For each enabled feature, what does it read? In plain words, written down. Live records, historic records, uploaded documents, whatever a person types, or nothing of yours at all. Most organisations can't answer this for at least one feature.

What of ours is retained, and for how long? Separate from what's read. Something can read a record without keeping anything, and something can keep every request indefinitely.

Does it respect the asker's permissions? Test it rather than asking. Have somebody with limited access ask a question whose honest answer would require information they can't see, and look at what comes back.

Where do people type free text? List those places. They're where the unexpected personal disclosures land, and they're usually governed differently from the structured records everybody worries about.

Could you answer an employee who asked? Somebody will, eventually, and the question will be simple: what does this thing know about me and who else can see it. If you can't answer plainly, that's the gap.

Run these five with whoever administers the system rather than whoever bought it. Administrators know which features are actually switched on, which is frequently not the same as the list anybody agreed to.

Six Things a Feature Might Be Reading, Reviewed

Your records inside the system at the moment of the request

It looks at current data to answer the question in front of it. It earns its place because this is what makes an answer specific to your organisation rather than generic, and specificity is usually the entire value.

Where it falls short is permissions and staleness. A system reading on somebody's behalf may not be bounded by what that person is allowed to see, and a record that's out of date produces a confident answer built on something no longer true.

Test the permissions question directly rather than accepting an assurance. Have somebody with limited access ask something they shouldn't get an answer to.

Staleness deserves its own check. A feature reading current records will happily answer from a record that stopped being accurate months ago, and nothing in the output indicates how old the underlying information is. Where an answer matters, knowing when the source was last touched matters with it.

The other thing worth knowing is what it does with a record that's incomplete. Missing information is extremely common in HR records, and a feature that fills a gap by inference rather than saying the gap exists is producing something quite different from what it appears to produce.

Your historic records accumulated over time

It draws on years of accumulated information rather than the current state. It earns its place where a pattern over time is the point, because a single snapshot can't show a trend.

Where it falls short is that history contains decisions rather than facts. Records reflect what your organisation did, including things it would now do differently, and anything drawing on them inherits that. There are also practical problems: definitions changed, systems changed, and the meaning of a field five years ago may not be its meaning now.

Useful for trends with somebody who knows the history alongside it. The fairness dimension of drawing on historic decisions is a governance question, covered elsewhere, and worth local advice.

The practical trap is a definitional one. Field meanings drift, categories get renamed and reused, and a system migration frequently mapped old values onto new ones in ways nobody documented. A trend computed across that boundary can look perfectly smooth and describe nothing real.

Worth asking how far back the feature actually reaches, as well. Some draw on everything available and some use a window, and the two produce noticeably different answers about anything that changed.

Documents people upload

Contracts, letters, policies, forms, whatever gets attached. It earns its place because documents hold information that was never structured, and extracting from them beats retyping.

Where it falls short is that documents contain far more than the thing you wanted. A contract carries personal details, an attached letter may reference a third party, and a scanned form may include information nobody intended to share. Upload also bypasses whatever controls sit on your structured records.

Worth establishing what happens to an uploaded file after the answer is produced. Whether it's retained, and where, is frequently different from how your records are handled.

The third-party problem is the one people miss. A document about one employee routinely names others: a manager, a colleague mentioned in a letter, somebody referenced in a complaint. Those people didn't upload anything and may have no idea their information is now in a system that processes it.

It's also worth deciding who can upload at all. The control on structured records is usually careful and the control on attaching a file is usually nothing, which is an odd asymmetry once you notice it.

Information the vendor holds from other customers

The feature draws on patterns or benchmarks derived from other organisations. It earns its place by offering context you can't generate alone, since your own data can't tell you what's unusual.

Where it falls short is opacity and reciprocity. You can't see whose information contributed, whether those organisations resemble yours, or how current it is, and the arrangement generally runs both ways, which means your material may be contributing to somebody else's version.

Ask what the comparison set is and whether you're in it. Both halves of that question deserve a plain answer.

Comparability is the substance of the answer. A benchmark drawn from organisations of a very different size, sector or country tells you about them rather than about you, and a figure presented without that context invites a conclusion nothing supports.

Currency matters too. A comparison set assembled some time ago and not refreshed describes a world that may have moved, and nothing about the way the number is presented will say so.

Information the underlying technology was built on

The capability was shaped by a large body of material long before your organisation existed to it. It earns its place because this is what allows a feature to handle language, documents and requests at all.

Where it falls short is that it's unknowable to you in any detail and unchangeable by you entirely. It's also the source of general-knowledge answers that sound authoritative about your specific situation while having no connection to it, which is a failure mode worth recognising.

Nothing to do here except know it's there. Where an answer sounds specific to you, ask what it read to produce it.

This is the source of a particular kind of confident wrongness worth recognising. A feature can produce a fluent, plausible statement about employment practice, entitlements or what organisations typically do, drawn entirely from general material and bearing no relation to your policies, your contracts or your jurisdiction.

The tell is that the answer contains no specifics from your organisation. If it could have been written for anybody, it probably was.

Information a person types into a prompt

Whatever somebody writes in the box, including everything they added as context. It earns its place because a person explaining their situation gets a better answer, and explanation means detail.

Where it falls short is that this is the least controlled channel you have. People type things they'd never put in a record: a colleague's circumstances, a health situation, a complaint in progress. It frequently sits outside the handling arrangements covering your structured data, and it's the route by which information leaves organisations unnoticed.

The one to establish properly before anybody uses a feature. Where prompts go, whether they're retained, and whether they train anything are three separate questions.

Worth telling people plainly what to avoid typing, too, in a sentence rather than a policy document. Most inappropriate disclosure into these boxes isn't carelessness, it's somebody trying to give enough context to get a useful answer, and a short piece of guidance at the point of use prevents most of it.

The same applies to information about third parties. Somebody describing a situation involving a colleague has put that colleague's circumstances into a system, and they generally haven't thought of it that way.

The Decision Table

Situation Scale Setup Primary Pain Recommended Starting Point
Both questions answerable in writing Any Any None Change nothing
Feature on, nobody knows what it reads Any Enabled Unknown exposure Write it down, per feature
Vendor answer was about training Any Evaluating Talking past each other Ask the three questions separately
Free text in use, unexamined Any Enabled Least controlled channel Establish where prompts go
Summarising across records Any Enabled Permissions may not carry Test with a limited account
Drawing on historic records Any Enabled History holds decisions Somebody who knows the history
Benchmarks from other customers Any Enabled Opaque, and reciprocal Ask what the set is, and if you're in it
Documents uploaded routinely Any Enabled More in them than intended Establish retention for uploads
Employee asks what it knows Any Any An answer you owe them Be able to answer plainly

The fifth row is the one that produces a concrete problem soonest, and it's testable in ten minutes. Give somebody with restricted access a question whose honest answer needs information they can't see, and look at what comes back. The result is either reassuring or immediately actionable.

The fourth row is the one that turns out to matter most and gets examined last. Structured records attract the governance attention because they're visible and countable. Free text attracts none, contains more sensitive material, and is where information actually leaves.

The ninth row is the one that sets the standard for all the others. If you can answer an employee plainly, in ordinary words, you almost certainly have the rest of this in order, and if you can't, the gap is somewhere in the rows above.

What Leaves and What Stays

Reading is not the same as taking. A feature consulting a record to produce an answer hasn't necessarily kept anything. Whether it did is a separate question with a separate answer, and conflating them produces both false alarm and false comfort.

Processing location is its own question. Where the computation happens determines whose systems your information passed through, which matters for reasons that differ by jurisdiction. Ask it plainly and take local advice on what the answer means for you.

Retention has a default and it's rarely zero. Requests are logged, outputs are stored, uploads are kept. All of that is reasonable and most of it is invisible unless you ask what's retained, for how long, and who can look at it.

Logs are records too. A log of who asked what, kept for perfectly sensible operational reasons, is itself a body of information about your people and their questions, and it rarely appears on anybody's list of what the system holds.

Improvement is the word to listen for. Language about the product learning or getting better with use frequently means your material contributes to it. That may be fine and it's a decision to make deliberately, with a written answer rather than an impression.

Prompts are usually governed separately. The arrangements covering your records often don't extend to what somebody typed, and prompts contain the most sensitive material you'll produce. This is the gap worth closing first.

The employee's question is the test. Somebody will ask what this thing knows about them and who can see it. If the answer requires a technical explanation, it isn't an answer. If you can't give one at all, you've found the work.

Third parties are in your records too. A document about one person routinely names others, and a prompt describing a situation usually involves a colleague. Those people made no choice about it, which is worth holding in mind when deciding what a feature may read.

The reason to be specific rather than anxious here is that vague concern produces either blanket prohibition or blanket permission, and both are worse than knowing. A feature reading current records, retaining nothing, and training nothing is a very different proposition from one where everything typed improves a shared product, and only one question separates them.

Where These Arrangements Go Wrong

The failure How it shows up What would have to change
Three questions answered as one Everybody agrees and disagrees Ask reading, training and retention separately
Permissions don't carry to the feature Somebody sees what they shouldn't Test with a limited account
Prompts outside the arrangements The sensitive channel, ungoverned Establish where typed text goes
Uploads retained by default Documents held nobody accounted for Ask about upload retention specifically
Historic records taken as fact Old decisions treated as evidence Somebody who knows what changed
No answer for an employee A simple question you can't meet Write the plain-words answer now

The second row is the one that produces an actual incident rather than a theoretical exposure. It's also the easiest to check on this entire list, and the check takes minutes, which makes it the first thing to do after reading this.

The sixth row is the one that arrives without warning. The question is always simple and always reasonable, it comes from somebody who has a right to ask, and the answer needs to exist before it's asked rather than be assembled while somebody waits.

The fifth row is the subtle one, because nothing about it looks like a failure. A trend drawn across records whose definitions changed produces a smooth, plausible line describing nothing that happened, and the only person who would catch it is somebody who remembers the change.

What to Put in Writing

Artefact Who owns it When it is written What it prevents
What each feature reads Whoever owns the system Now Unknown exposure
What's retained, and for how long Whoever owns the system Before enabling Retention nobody accounted for
Whether your material trains anything Get it from the vendor, in writing Before enabling A verbal assurance you can't rely on
Where prompts and uploads go Whoever owns the system Before anybody uses it The ungoverned channel
The permissions test result Whoever owns the system Before wide rollout Somebody seeing what they shouldn't
What applies to you, per jurisdiction You, with local advice Before enabling A position you assumed

The third row is the one to insist on having in writing rather than in a conversation. It's the question employees and auditors actually ask, the answer may change at a vendor's discretion, and a sales assurance given a year ago is not something you can produce when somebody wants to see it.

The fourth row is the one to write before a rollout rather than after one. Once people have started using a feature, the prompts already exist, and whatever they contain has already gone wherever it goes.

Questions to Ask Before You Commit

On reading. What does it consult to answer a question? A bad answer is your data.

On training. Does anything we put in improve this for others? A bad answer is that it's anonymised.

On retention. What's kept after the answer, and for how long? A bad answer is only what's necessary.

On prompts. Where does typed text go? A bad answer is that it's treated the same way.

On permissions. Does it respect what the asker can see? A bad answer is that access is controlled.

On location. Where does the processing happen? A bad answer is in the cloud.

What Getting This Wrong Costs

The first cost is an answer you owe somebody and can't give. An employee asking what a system knows about them, and who can see it, is asking a reasonable question about their own information. Not being able to answer plainly is a trust problem on its own, separate from whatever the actual arrangement turns out to be, and assembling the answer while they wait makes it worse.

The unease also spreads faster than the answer does. One person asking and getting a vague response tells colleagues, and by the time a proper answer exists it's arriving into a conversation that has already formed its own view.

The second cost is the disclosure nobody designed. Free text is where people explain themselves, explanation means detail about their circumstances and frequently about somebody else's, and prompt handling routinely sits outside the arrangements covering everything else. Organisations spend their attention on structured records because those are visible and countable, while the genuinely sensitive material goes out through a box nobody governs.

The asymmetry is worth stating plainly, because once seen it's hard to unsee: adding a field to a record involves a review, and typing a paragraph about somebody's health into a prompt involves nothing at all.

The third cost is deciding by accident. A feature that reads current records and retains nothing is a different proposition from one where everything contributed improves a shared product, and the difference is a single question asked once. Organisations that never ask it don't end up with the safer arrangement by default; they end up with whatever the vendor's default is, having made no decision at all.

Defaults also move. An arrangement that was one thing when you enabled a feature can be another thing a year later, changed in a release note nobody read, which is why the answer is worth having in writing and worth revisiting rather than settling once.

So do three things this week. Write down what each enabled feature reads. Ask the training question in plain words and get the answer in writing. And run the permissions test, because it takes ten minutes and it's the one that produces an incident.

When You Are Ready to Go Further

Start with the inventory, feature by feature: what it reads, what's retained, whether anything of yours improves the product. Three columns, an afternoon's work, and the exercise reliably turns up at least one feature where nobody can fill in a box. That blank is the finding.

Do it from the system's own configuration rather than from a list of what was purchased. The two diverge, because features arrive enabled in releases, and the configuration is the one that describes what's actually reading your records.

Then close the prompt gap, because it's the widest one and the least examined. Establish where typed text goes, whether it's kept, and whether it trains anything, and make sure whatever governs your records extends to it. People will type things into that box that they'd never enter into a record, and they'll do it because they're trying to get a useful answer.

A short line of guidance at the point of use does most of the work here. Telling people plainly what not to type, where they're about to type it, prevents far more than a policy nobody opens, and it takes one sentence.

Finally, write the plain-words answer to the employee question before somebody asks it. What the system reads, what it keeps, who can see it, in language that needs no technical background. If you can't write that, you don't yet know your own arrangement, and that's worth discovering now rather than in the conversation itself.

Share it before anybody has to ask, if you can. A short, plain description of what a new feature reads, published when it's switched on, removes most of the unease that otherwise accumulates and reaches you later as a complaint.

HROpsLab publishes independent comparison work across HR tooling. We sell nothing, we take no vendor money, and we publish no paid placements. If the next step is understanding what your current tooling does here, our comparison work is one place to start.


Frequently Asked Questions

What data do AI features in HR tools actually use?

It depends on the feature, and the honest answer usually involves several sources at once: your current records, accumulated historic records, documents people upload, whatever somebody types into a prompt, and the material the underlying capability was built on long before it encountered your organisation. Some features also draw on information derived from other customers. The useful exercise is to write this down per feature rather than treating it as one topic, because the answers differ and so do the consequences, and most organisations find at least one feature where nobody can fill in the box.

Is your employee data training the vendor's model?

This is the question most people mean and it's worth asking in exactly those words, then getting the answer in writing. It may legitimately be no, it may be yes with an option to decline, or it may vary by feature. What you shouldn't accept is an answer about anonymisation or security, which addresses something different. A verbal assurance in a sales conversation isn't something you can produce a year later when an employee or an auditor asks, and vendor defaults in this area can change.

What's the difference between training data and what a feature reads?

Training material is what shaped the capability before it ever met your organisation, and you generally can't see it or change it. What a feature reads is the information it consults to answer the particular question in front of it, which may be your live records, your history, an uploaded document, or nothing of yours at all. The two are independent: a feature can read your records without your records training anything, and it can be built on vast material while touching nothing of yours at the moment of use.

What happens to information typed into an AI prompt?

Establish this before anybody uses the feature, because it's the least governed channel most organisations have. People explain their situation when they ask a question, and explanation means detail about their own circumstances and frequently about a colleague's. Prompt handling is often covered by different arrangements from your structured records, with different retention and different treatment, and the gap goes unexamined because prompts aren't visible or countable the way records are. Ask where they go, what's kept, and whether they train anything.

Should AI features be allowed to read historic HR records?

They can be, and the thing to hold onto is that historic records contain decisions rather than facts. They reflect what your organisation did, including things it would now do differently, and anything drawing on them inherits those patterns. There are practical issues too: definitions shifted, systems changed, and a field's meaning some years back may not be its meaning today. Keep somebody who knows that history alongside any output drawn from it. The fairness dimension is a governance question covered elsewhere, and worth local advice.

Can an AI feature see more than the person asking it?

Frequently yes, and this is a common gap rather than an exotic one. Permissions applied to a person don't automatically constrain a system reading records on their behalf, and a summary drawn across your data may include things the asker has no right to see, with nothing in the output indicating it. It's also the easiest thing on this list to check: have somebody with restricted access ask a question whose honest answer would require information they can't see, and look at what comes back.

Does your data leave your organisation when you use an AI feature?

That depends on where the processing happens and what's retained afterwards, and those are two separate questions that get answered as one. Something can read a record without keeping anything, and something can retain every request indefinitely. Where the computation occurs determines whose systems your information passed through, which matters for reasons that differ by jurisdiction. Ask both plainly, get the answers in writing, and take local advice on what they mean for your obligations.

What should you establish before switching on an AI feature?

Three things, and each takes minutes. What the feature reads to produce an answer. What's retained and for how long. Whether anything you supply improves the product for anyone else. Then run the permissions test, because it's quick and it's the one most likely to surface a concrete problem. Separately, establish what applies to you with local advice, since obligations around personal data differ by jurisdiction and sector and change, and the arrangement you're comfortable with may not be the one you're permitted.

Reading is one question. Keeping is another. Improving is the one people mean.

Share on X Share on LinkedIn

What to do next?

Explore More Articles

Dig deeper into HR Ops strategy, tools, and workflows built for real teams.

Browse the blog →
Join the HROpsLab Community

Connect with People Ops practitioners sharing real workflows, tools, and challenges.

Join now →