TL;DR
- The core decision: what you'd do if an AI feature started behaving differently without warning.
- When doing nothing is right: when nothing downstream depends on the behaviour staying put.
- What has to be true: you'd notice, and you could say roughly when it started.
- How the options split: by whether the change is theirs, yours, or nobody's decision at all.
- Decision rule: keep something fixed to compare against, because nothing else tells you.
- Outcome to expect: the same feature, and a date you can put on any change in it.
Nothing Changed on Our Side
Somebody says the summaries have got shorter. Or that the categories are landing differently. Or that a figure that used to come through reliably now doesn't.
You check. No release on your side, no configuration change, nobody touched the settings. The obvious conclusion is that somebody is misremembering, because people misremember all the time and there's nothing to point at. So it gets dropped, and a month later two more people say something similar.
This is an ordinary problem with a specific shape. Software you buy changes on somebody else's schedule, which has always been true. What's different when the changing part is a model is that the change doesn't have to appear in a release note, doesn't have to alter any interface, and produces no error. The feature does what it always did, and produces something slightly different, and nothing marks the transition.
Best tools for AI HR Tools
The reframe that makes this manageable: this isn't a technical problem, it's a supplier one. You don't need to understand what changed. You need to know what can change, whether you're told, whether you can stay where you are, and how you'd notice. Those are questions you ask a vendor and questions you arrange your own process around, and none of them require knowing anything about how the thing works.
The awkward part is that most organisations have no way to answer the last one. Without something fixed to compare against, there's no moment when anybody can say the behaviour is different now, and the conversation stays at the level of somebody thinking the summaries got shorter.
Where an output informs a decision about a person, and the behaviour producing it moves, questions about what you're obliged to assess, record or disclose may follow. Those differ by jurisdiction and are changing. Establish yours with local advice, and take the governance framing from the AI in the workplace material.
When You Genuinely Do Not Need to Act Yet
Nothing depends on the behaviour staying put. The output is a draft somebody reworks, and variation between one month and the next costs nothing. Then this is a non-problem and worth recognising as one.
Somebody has said results feel different. Worth taking seriously rather than dismissing, because it's the only signal most organisations have. It's also usually unprovable, which is the thing to fix.
A process was built on the old behaviour. Somebody downstream expects a particular shape, length or format, and a change breaks something. At that point behaviour stability is a real dependency.
The edge case that forces it. An error surfaced and nobody can say whether the feature has always done this or started recently. That's the position worth never being in again, and the fix is cheap.
Five Questions This Reader Asks at 11pm
Why is it answering differently? Several possible reasons and they're worth separating. The vendor changed the underlying capability. The vendor changed how their product uses it. The feature adapts as data accumulates. Your own data shifted. A default changed in a release. Or nothing changed and the question was phrased differently this time, which is more common than people expect.
Will they tell us? Sometimes, and it depends on what you agreed and on how the vendor treats this class of change. A product update usually gets a release note; a change in the underlying capability frequently doesn't, because from the vendor's perspective nothing about their product changed. That distinction is worth asking about explicitly.
Can we stay on what we've got? Occasionally, and it's worth asking even when the answer is no, because the answer tells you how the vendor thinks about this. Pinning behaviour is more common in products built for this kind of dependency and rare in features bundled into general software.
How would we know? Only by comparing against something fixed. There's no notification for this, no log entry and no error. A small set of retained examples with their original outputs, rerun occasionally, is the whole method.
Isn't it improving? Possibly, and improvement in general isn't the same as improvement for you. A change that raises overall quality can be worse on your particular cases, and the direction for you is unknowable without checking your own examples.
What Can Change Without You
| The change | Would you be told | What it does to a process built on the old behaviour |
|---|---|---|
| Underlying capability updated | Frequently not | Outputs shift in shape, tone or emphasis |
| How the product uses it | Sometimes, in a release note | Behaviour changes with no visible cause |
| Feature adapts to accumulated data | Almost never | Gradual drift, no single moment |
| Your own data shifts | No, this one is yours | Same feature, different results |
| A default changes in a release | Usually, if you read them | Settings you thought were yours |
| Capability withdrawn or replaced | Yes, eventually | Something downstream stops working |
| Thresholds retuned | Rarely | More or fewer items flagged, silently |
| A source it draws on changes | No | Answers move with no product change |
The first row is the one people mean and the one hardest to get a straight answer about. From a vendor's perspective, swapping what sits underneath their feature is an internal implementation decision, and many genuinely don't consider it a customer-facing change. Whether you agree is worth establishing before it happens.
The third row is the one that defeats detection entirely. A gradual shift has no transition to notice, so there's never a day when anybody says it's different now, and by the time the difference is obvious nobody can date it. This is the strongest argument for a fixed comparison set.
The fourth row is worth separating out because it's yours rather than theirs. A feature behaving differently after you restructured, changed a policy or onboarded a new population hasn't changed at all. That's an important distinction when you're deciding who to talk to.
Five Diagnostic Questions You Can Self-Assess Against
What would tell you the behaviour moved? If the honest answer is somebody mentioning it, you have no detection. That's the common position and the one to fix first.
Do you keep any fixed examples? A handful of real inputs with their outputs recorded. Almost nobody does this and it's the entire method for answering every other question here.
What downstream depends on the shape of the output? Length, format, the presence of a particular field. Anything that does is a place where a change becomes a break rather than a difference.
Do you read the release notes? Genuinely, not in principle. Default changes usually are announced, and they're usually announced somewhere nobody looks.
Could you tell their change from your change? If your organisation restructured last month and the outputs shifted, those are two candidate explanations and you need a way to tell them apart.
Run these five with whoever uses the output daily rather than whoever administers the system. Users notice behaviour shifts first, months before anything reaches an administrator, and they usually mention it to a colleague rather than reporting it.
Six Ways the Behaviour Moves, Reviewed
The vendor updates the underlying capability
What sits beneath the feature is swapped or upgraded. It earns a place first because it's the change most likely to alter behaviour noticeably and least likely to be announced.
Where it causes trouble is the absence of any marker. No release, no error, nothing in the interface. Outputs shift in length, emphasis or format, and the only evidence is people saying things feel different, which is easy to dismiss.
It's easy to dismiss for a good reason, too: people do misremember, and most reports of this kind turn out to be nothing. That's precisely why the comparison set matters. It lets you take the report seriously at no cost, instead of choosing between believing an impression and ignoring one.
Ask directly whether this counts as a notifiable change for them. The answer, whatever it is, tells you what to expect and whether you need your own detection.
The reason a vendor may genuinely not consider it notifiable is worth understanding rather than resenting. From where they sit, they've kept the feature doing what it always did and improved what's underneath. From where you sit, the outputs your process depends on now look different. Both descriptions are accurate.
What you're really asking is whether behaviour is part of what you bought or whether only the capability is. Vendors differ on that and very few state it unprompted, so the question is worth putting plainly and writing down the answer.
The vendor changes how the product uses it
The capability stays and the way the product employs it changes: different instructions, different context, different post-processing. It earns its place because it's common and because it's genuinely their product changing.
Where it causes trouble is that these changes are frequently considered improvements and shipped quietly. The vendor is right that they improved something, and your particular use may have depended on the previous behaviour.
This is also the category where raising it gets the best response, because it's unambiguously their product and they can see what they changed. A specific report with examples frequently produces an explanation and sometimes an option to restore the previous behaviour.
More likely to appear in release notes than the first case. Worth reading them for this specifically rather than scanning for new features.
The phrasing works against you here. Release notes describe what was added and improved, and a change to how an existing feature behaves gets written up as an enhancement rather than as a difference. Reading for the word improved, in connection with something you rely on, catches more than reading for the word changed.
It also helps to know which of your features are actually used heavily. Scanning notes is much faster when you're looking for four specific things rather than reading everything.
The feature adapts to accumulated data
Behaviour shifts as more of your information builds up. It earns its place when the adaptation is the point, because a feature that fits your organisation better over time is doing something valuable.
Where it causes trouble is that there's no moment of change to detect. The drift is continuous, nobody can date it, and comparing against what you remember is unreliable because memory of outputs is poor.
It's worth asking what the adaptation is fitted to, as well. A feature learning from what people accepted is learning from approvals, and if those approvals were nominal, it's fitting itself to whatever passed rather than to whatever was right.
Only detectable with fixed examples rerun periodically. Ask whether the feature adapts at all, because it isn't always obvious and it changes what detection you need.
Adaptation also has a direction you didn't choose. A feature fitting itself to your accumulated data is fitting itself to what your organisation has done, including the parts you were trying to move away from, and nothing about it will indicate that.
Where it adapts, the comparison set needs rerunning more often than where it doesn't, because there's no event to prompt a check and the change accumulates quietly between reruns.
Your own data shifts
The feature is unchanged and your information isn't. A restructure, a new population, a policy change, a different mix of cases. It earns its place because it's the explanation people forget, and it's frequently the right one.
Where it causes trouble is misattribution. A team convinced the vendor changed something can spend a long time on that conversation when the cause is a reorganisation nobody connected to it.
The connection is genuinely hard to see from inside, because the two things happen in different parts of the organisation. Whoever restructured isn't thinking about summarisation outputs, and whoever reads those outputs may not know a restructure happened.
Check your own timeline first. It's free, it's fast, and it resolves a decent proportion of these.
Keeping a light change log makes this trivial rather than archaeological. A dated line whenever something structural happens, a restructure, a policy update, a new population onboarded, is enough to answer the question in a minute instead of an afternoon.
The cases that catch people are the ones where both things happened. Your data shifted and the vendor changed something, in the same quarter, and untangling which produced which effect needs the comparison set rather than the timeline.
A default changes in a release
A setting you never chose moves, or a new option arrives switched on. It earns its place as the most tractable item here, because it's usually documented and usually reversible.
Where it causes trouble is that release notes are long, read by few people, and phrased in terms of improvements rather than changes to existing behaviour. The information was available and nobody saw it.
That also makes it the easiest category to fix completely. One person, a short list of features you actually rely on, and a few minutes after each release closes almost all of it.
Worth assigning to somebody specific rather than to everybody. A person who reads release notes for behaviour changes catches most of this category.
It takes very little time when it's somebody's named job and it never happens when it's everybody's. The failure here is organisational rather than informational: the notice was published, it was accurate, and it went to a distribution list where reading it was nobody's particular responsibility.
Worth pairing with a quick check of anything you rely on after a significant release, rather than waiting for the scheduled comparison.
The capability is withdrawn or replaced
A feature goes away, or is superseded by something that works differently. It earns its place because it's the one change you'll definitely be told about.
Where it causes trouble is the notice period and whatever you built on it. Anything downstream that depends on the output has to be reworked on somebody else's timeline, and that timeline is rarely generous.
Replacement is the harder version of withdrawal, because the feature is still there and behaves differently. Nothing stops working, everything shifts, and the transition is easy to underestimate precisely because it doesn't look like a removal.
The argument for knowing what depends on each feature before you need to know. That list takes an afternoon and is impossible to produce quickly under pressure.
The dependencies that hurt are rarely the obvious ones. A report somebody built that assumes a field, a downstream step expecting a particular length, a spreadsheet a team maintains from the output: these accumulate without anybody recording them and surface only when the output stops arriving in the expected shape.
Ask the people who use the output rather than the people who own the system. The users know what they've built on it, and much of it won't be documented anywhere.
The Decision Table
| Situation | Scale | Setup | Primary Pain | Recommended Starting Point |
|---|---|---|---|---|
| Nothing depends on stable behaviour | Any | Any | None | Change nothing |
| People say results feel different | Any | Any | Unprovable either way | Start keeping fixed examples |
| Nothing fixed to compare against | Any | Any | No detection at all | A handful of examples, today |
| A process expects a particular shape | Any | Any | Change becomes a break | List what depends on the output |
| Organisation changed recently | Any | Any | Two candidate explanations | Check your own timeline first |
| Release notes unread | Any | Any | Announced changes missed | Assign them to a person |
| Feature adapts over time | Any | Adaptive | No moment to detect | Fixed examples, rerun on a schedule |
| Vendor won't say what can change | Any | Evaluating | Unknown exposure | Weigh it as evidence |
| Feature being withdrawn | Any | Any | Somebody else's timeline | Know what depends on it already |
The third row is the one to act on today, because it's the precondition for every other row and it takes an hour. A handful of real inputs, their outputs recorded, stored somewhere outside the system. That single artefact converts every future version of this question from unanswerable into a morning's work.
The fifth row is the cheapest diagnosis available and the most frequently skipped. Before opening a conversation with a vendor about changed behaviour, check what changed on your side. A surprising proportion of these resolve there.
The eighth row is the one to weigh at purchase rather than later. A vendor who won't say what can change is telling you that behaviour stability isn't part of what they're selling, which may be perfectly acceptable and should be a decision rather than a discovery.
Noticing That Results Moved
Memory of outputs is unreliable. People genuinely cannot recall what a summary looked like two months ago, which is why these conversations stall. The disagreement isn't about facts, it's that nobody has any.
And the more confident the recollection, the less it should settle. Somebody certain that outputs used to be better is describing an impression formed over time, which is exactly the thing memory is worst at.
A fixed set is the whole method. Real inputs, their outputs recorded with a date, stored outside the system. Rerun them occasionally and compare. There's no more sophisticated technique available and no simpler one needed.
A dozen is enough. The temptation is to build something comprehensive, which makes the rerun expensive and therefore skipped. A small set that actually gets rerun beats a thorough one that doesn't.
Keep them somewhere the system can't touch. A baseline held inside the thing it's meant to check moves when that thing moves. A file elsewhere is fine and is the point.
Include your awkward cases. Ordinary inputs are stable across changes more often than edge cases are. If you want early warning, the examples that stress the feature are the ones that will show a difference first.
Keep the inputs, not only the outputs. A recorded result with no record of exactly what was asked can't be rerun, which quietly makes the whole set useless at the moment you need it.
Rerun on a schedule, not on suspicion. Running the comparison only when somebody complains means you find out after the complaint. A quarterly rerun costs an hour and catches things before they reach anybody.
Expect some variation that means nothing. Not every difference is worth chasing, and a set that fires on trivial wording changes trains people to ignore it. What you're looking for is a shift in substance, length or structure showing up across several examples at once.
Record what depends on each output. When a capability is withdrawn or changed materially, the first question is what breaks, and assembling that answer under time pressure is far harder than writing it down once.
Note the version or date alongside each result. A comparison is only useful if you can say what you're comparing against, and a folder of outputs with no dates is very nearly as unhelpful as no folder at all.
The reason this is worth the small effort is that it converts an argument into a comparison. Somebody says the outputs changed, you rerun twelve examples, and either they did or they didn't. That takes an hour and it ends a discussion that would otherwise run for weeks and conclude nothing.
It also changes how people report things. Once colleagues know there's a way to check, they mention differences earlier and more precisely, because raising it leads somewhere instead of into an unwinnable argument about memory.
Where These Arrangements Go Wrong
| The failure | How it shows up | What would have to change |
|---|---|---|
| No baseline kept | An unprovable argument about memory | A dozen fixed examples, dated |
| Baseline held inside the system | It moves when the thing moves | Store it elsewhere |
| Only ordinary cases retained | Change shows up late | Include the awkward inputs |
| Rerun only on complaint | Detection after the fact | A scheduled rerun |
| Dependencies never listed | Scramble when something is withdrawn | Write the list once |
| Own changes not considered | Vendor conversation about your restructure | Check your timeline first |
The first row is the root of everything else here. Without a baseline, every question in this piece reduces to people disagreeing about what they remember, and that disagreement can't be settled by anybody.
The sixth row is the avoidable embarrassment. Opening a support case about changed behaviour, and discovering partway through that the change was a restructure on your own side, costs credibility you'll want later when the change genuinely is theirs.
The third row is the one that delays detection without anybody doing anything wrong. A comparison set built only from clean, typical inputs will keep matching for a long time after the behaviour has started moving on the cases that actually stress the feature.
What to Put in Writing
| Artefact | Who owns it | When it is written | What it prevents |
|---|---|---|---|
| A dozen fixed examples and their outputs | Whoever owns the system | Today | An unanswerable question later |
| The rerun schedule, and who does it | Whoever owns the system | With the examples | Detection only after a complaint |
| What depends on each output | Whoever owns the process | Now | A scramble on somebody's timeline |
| What the vendor says can change | Whoever evaluates | Before committing | Assuming stability nobody promised |
| Whether you'll be notified, and of what | Whoever evaluates | Before committing | Silent changes you agreed to |
| Your own change log | Whoever owns the process | Ongoing | Blaming a vendor for your restructure |
The first row is the item to do before finishing this article. A dozen real inputs, their current outputs, today's date, saved somewhere outside the system. It costs an hour, it needs nobody's approval, and it's the only thing that makes the rest of this answerable.
The sixth row is the one that pays for itself the first time it's used. A dated line whenever something structural changes on your side turns the most common false alarm in this area into a thirty-second check.
Questions to Ask Before You Commit
On notification. What changes will you tell us about? A bad answer is that you publish release notes.
On the underlying capability. Does swapping it count as a change? A bad answer is that it's an implementation detail.
On pinning. Can we stay on current behaviour? A bad answer is that everybody gets improvements.
On adaptation. Does the feature change with use? A bad answer is that it learns.
On comparison. Can we retain examples and rerun them? A bad answer is that outputs aren't deterministic.
On withdrawal. What notice do we get? A bad answer is that it's in the terms.
What Getting This Wrong Costs
The first cost is an argument nobody can settle. People report that results feel different, there's no evidence either way, and the discussion runs for weeks and concludes with everybody holding their original position. That consumes real time, it erodes confidence in the feature, and it's entirely preventable by an hour's work done in advance.
The second cost is a change you can't date. When something is eventually confirmed as different, the immediate question is how long it's been that way, because everything produced in the interval is now uncertain. Without a baseline the honest answer is that nobody knows, which turns a bounded problem into an unbounded one and makes any remedy either enormous or arbitrary.
It also makes the vendor conversation much weaker. Raising a concern with a date and a dozen before-and-after examples is a different exchange from raising one with an impression, and the two get very different responses.
The third cost is discovering a dependency at the worst moment. A capability is withdrawn or materially changed, and only then does anybody work out what was built on it: the report that expects a particular format, the process that assumes a field, the downstream step that breaks. Assembling that list under a vendor's timeline is far harder than writing it once when nothing is urgent.
The list also tells you something useful long before any withdrawal. A feature with a dozen things hanging off it is a different risk from one nothing depends on, and knowing which is which changes how much detection each deserves.
So do three things. Save a dozen real examples with their outputs today. Put a rerun in the calendar with a name against it. And write down what depends on each feature's output, because that's the list you'll be asked for on the day you least want to build it.
All three sit entirely on your side of the relationship, which is what makes them worth doing regardless of what any vendor tells you. They also work the same way for every feature you adopt afterwards.
When You Are Ready to Go Further
Start with the examples, and start now rather than after the next conversation about this. A dozen real inputs covering your ordinary cases and your awkward ones, with the outputs exactly as they come back today, stored outside the system in something plain. That artefact is the whole of the detection method and it costs an hour.
Store them somewhere a colleague could find without you. A comparison set living in one person's folder disappears when they change role, which is usually around the time somebody starts asking whether behaviour has moved.
Then put the rerun on a schedule with a person's name attached. Quarterly is enough for most situations. What matters is that it happens without anybody complaining first, because detection triggered by a complaint always arrives after whatever the complaint was about.
Run it after any significant release as well, not only on the schedule. Those are the moments when something is most likely to have moved, and the comparison takes under an hour once the examples exist.
Finally, ask the notification question of your vendors, and write down the answer. Whether they'd tell you if the underlying capability changed, what counts as notifiable, and whether you could stay where you are. Even a disappointing answer is useful, because it tells you exactly how much of the detection has to be yours.
Ask it of the vendors you already use, not only the ones you're evaluating. The features already running in your organisation are the ones with processes built on them, and nobody asked this question when they were switched on.
HROpsLab publishes independent comparison work across HR tooling. We sell nothing, we take no vendor money, and we publish no paid placements. If the next step is understanding how the tools in this space handle change, our comparison work is one place to start.
Frequently Asked Questions
Why does an AI feature give different answers to the same question?
Several reasons worth separating, because they lead to different actions. The vendor may have updated what sits underneath the feature. They may have changed how their product uses it. The feature may adapt as your data accumulates. Your own data may have shifted, which is the explanation people forget and frequently the right one. A default may have changed in a release. Or nothing changed and the question was phrased slightly differently, which happens more often than anybody expects.
Do vendors tell you when the underlying model changes?
Sometimes, and the honest answer is that it depends on how they classify the change. A product update generally gets a release note. Swapping what sits beneath a feature frequently doesn't, because from the vendor's perspective their product didn't change, only its implementation. That distinction is worth raising explicitly before you commit, since the answer tells you whether you can rely on being told or whether all the detection has to be yours.
What does model drift mean in practice?
Behaviour moving over time without anybody deciding it should. It can come from the vendor's side, from a feature adapting as data accumulates, or from your own information changing so that identical behaviour produces different results. What makes it awkward rather than merely annoying is the absence of a transition to notice: gradual change has no moment, so nobody can point at a day, and by the time a difference is obvious it's impossible to date.
Can you stay on an older version of an AI feature?
Occasionally, and it's worth asking even where you expect the answer to be no, because the response tells you how the vendor thinks about behaviour stability. Products built for customers with this kind of dependency are more likely to offer it. Features bundled into general-purpose software rarely do, since the vendor's model is that everybody receives improvements. Where pinning isn't available, your own comparison set becomes the only instrument you have.
How do you tell if AI results have changed?
Keep a dozen real inputs with their outputs recorded and dated, stored somewhere outside the system, and rerun them periodically. That's the whole method and there isn't a more sophisticated one that's practical. Without it, the conversation reduces to people disagreeing about what they remember, and memory of outputs is genuinely poor. Include your awkward cases as well as ordinary ones, because edge cases tend to shift first and give you earlier warning.
What should you ask a vendor about changes?
Four things. What changes they'd notify you about, and specifically whether swapping the underlying capability counts. Whether you can remain on current behaviour if you need to. Whether the feature adapts with use, since that changes what detection you need. And what notice you'd get if a capability were withdrawn. Write down the answers, because these are exactly the questions that matter a year later and exactly the ones nobody can reconstruct from memory.
Do AI features get better over time?
They change, and better is doing a lot of work in that sentence. A change that raises overall quality across a vendor's whole customer base can be worse on your particular cases, especially if your data or your use is unusual. The direction of travel for you specifically is unknowable without checking your own examples, which is another argument for keeping a comparison set. Improvement claims are about aggregate behaviour and your experience is not aggregate.
What happens when an AI capability is withdrawn?
You'll be told, which makes this the one change with reliable notice, and then you're working to somebody else's timeline. The difficulty is rarely the feature itself and usually what was built on it: a report expecting a particular format, a process assuming a field exists, a downstream step that silently depends on the output's shape. Assembling that dependency list while a deadline runs is considerably harder than writing it down once, which is why it belongs on the list of things to do before you need it.
Nothing changed on your side. That doesn't mean nothing changed.