TL;DR
- The duty is yours: The employer running the tool owes the audit, the publication, and the notice. The vendor isn't a substitute.
- Do nothing only if: You don't screen or rank candidates or employees in NYC with an Automated Employment Decision Tool. Everything else is worth checking properly.
- What has to be true: A genuinely independent auditor ran the bias audit on the tool as deployed for your roles, the summary of results is published, and candidates got the required notice. The timing and wording of that notice sit in the city rules, so read the current version with counsel rather than from memory.
- How the options split: In-house audit teams, vendor-published audits, specialist AI audit firms, statistician academics, consultancy risk practices, or employment law firms coordinating a subcontracted auditor. Each has a real weakness.
- The test to apply: Has the auditor been involved with this tool before, and does the audit describe your applicant pool or the vendor's whole customer base? Those two answers are where exposure concentrates. Take them to local counsel rather than settling them yourself.
- What to expect: Civil penalties enforced by the NYC Department of Consumer and Worker Protection run up to 500 dollars for a first violation, and between 500 and 1,500 dollars for each subsequent one. The bigger cost is the one that doesn't appear on the invoice.
The Tuesday the Vendor Said It Was Handled
Priya runs talent operations for a 400-person healthcare group. She screens every nursing candidate with a video interview tool that scores responses and flags strong fits. On a Tuesday in late August, her compliance lead forwards a question from a candidate about NYC Local Law 144. The candidate wants to see the bias audit. Priya writes to the vendor. The vendor sends a PDF labelled "Annual Bias Audit". It looks official. It cites the right statute. It doesn't name Priya's roles or her applicant pool. But it's a vendor audit, and that's the first problem.
Priya's instinct is that the vendor has handled it. Her chief people officer is asking whether the company is exposed. Her CIO is asking whether the audit needs to be re-run. The board wants a one-line answer before the next quarterly review. And the candidate is still waiting for a response.
The real issue isn't whether the vendor has performed an audit at all. The real issue is whether Priya's company has performed the audit, on the tool as it actually runs for her roles, by someone who never touched the tool before, with the results published and the candidates notified. The vendor's audit answers a different question than the one the law asks.
Best tools for AI in the Workplace
When You Can Skip This Whole Piece
Four honest stages, and only the first lets you off the hook.
Stage one: you're out of scope. You don't use an Automated Employment Decision Tool to screen or rank candidates or employees for any role based in New York City. Picture a distribution business in Newark with no NYC positions and a hiring process that runs out of a shared inbox. Nothing scores anyone. Two things have to be true here, and both are worth checking rather than assuming. First, no position you post is based in New York City, which is a question about the requisition and not about where the candidate lives. Second, nothing between application and offer changes who a human ends up seeing. The ATS is the one everybody remembers. The sourcing tool that ranks inbound CVs by fit, before anyone opens the queue, is the one that gets missed. That tool screens. Nobody bought it thinking so.
Stage two: low friction. You use a tool, it has been live for more than a year, the vendor publishes a per-customer audit, and your applicant pool is large enough for the numbers to mean something. Your candidates already see the notice inside the application flow. Think of a Manhattan agency hiring for one role family at steady volume. Your job here is a diary entry, not a panic project. What actually breaks at stage two is the website. The audit gets refreshed on schedule, the summary page doesn't, and for months you're publishing a result that no longer describes the tool you run. Put a named owner on that page. Not a team. A person.
Stage three: real risk. You screen candidates in NYC, your vendor handed you a generic audit, and nobody can tell you whether candidates saw any notice at all. You have weeks, not quarters. The instinct is to procure the audit first, because that's the expensive, visible piece with a budget line attached. Do the notice first. Finding and engaging an independent auditor takes time you can't compress, while a missing notice is failing quietly on every application submitted today. Go and look at where yours actually sits. If it's in a careers-page footer, under the cookie banner, on a page candidates only reach after they've already applied, ask yourself honestly whether anyone has seen it.
Stage four: the edge case. Your tool is configured differently for different roles, or you switched vendors part-way through the year, or your applicant pool is small enough that the result comes back underpowered. A hospital hiring perfusionists might see eleven applicants across a year. An auditor asked to compare outcomes across categories in that pool has to say something confident about groups of three or four people, which no test statistic can honestly do. There isn't a clean answer. What you control is how the scope is written and how plainly the small counts get described in what you publish. Take local advice before you decide on that wording.
Five Questions You Are Asking at 11pm
Does the vendor audit count? No. The obligation sits with the employer running the tool, not the company that built it. A vendor audit measures the model across a pooled population that includes industries you don't hire in. Your applicant pool has its own shape, and that shape is what the law is interested in. Get this wrong and the sequence is predictable. The PDF goes in a folder. Someone tells the board it's handled, and the next person to read it carefully is a candidate's lawyer. You then commission the real audit under time pressure, with a complaint open and a deadline set by somebody else.
What does "independent" actually mean? The auditor can't have been involved in using, developing or distributing the tool. Past work for you on unrelated systems is fine. Past work for the vendor on this tool isn't. The conflict runs to the tool, not to you, and that's the part people invert. Ask for the relationships in writing, and ask about referral fees as well as project history, because money moves between firms that have never shared a file. Get it wrong and you've bought something that falls over the first time anyone asks who paid whom. Then you pay again.
What if the audit finds bias? You publish the results anyway. The audit is a measurement, not a pass mark. A finding also hands you something to act on: a cut point worth revisiting, or a conversation with the vendor about what the model rewards. The real hazard is the quiet re-run, where someone asks the auditor for different parameters until the numbers look better. Every one of those requests sits in an email thread with a timestamp. Concealment reads far worse in a deposition than a disparity you found and started fixing.
Does this apply to remote roles? If the role is based in New York City, the law applies to the screening of candidates for that role, regardless of where the candidate sits. The anchor is the position, not the person. Think of a fully remote engineering role carrying an NYC office address for payroll reasons. It's an NYC position with a national applicant pool, often the largest pool in the company. Teams that scope their notice by candidate location miss exactly that group. Confirm the facts of your own case with local counsel before you write anything into a policy.
How much does an audit cost? The fee moves with the size of your applicant pool and the number of role profiles you want covered. Don't budget off a vendor blog post. Ask for two written quotes and read the scope before you read the number. One may price a single pooled test across everything you run. Another may price each role profile separately, with an explicit check on whether your counts support a finding. Those are different products at similar headline prices. Buy the cheap one and you'll discover in month four that it covered one role, then buy the second anyway.
The Three Categories This Splits Into
Specialist AI audit firms. Firms whose full-time work is bias auditing for employment tools. They know the test statistics and the shape of the summary the city expects to see, which means you aren't paying someone to learn the statute on your clock. You can spot a real one at the intake stage: the questionnaire asks for your role profiles by name and your applicant counts for each, not your headcount and your logo. They're the right call when your pool is large and the result has to survive a reading by someone hostile. Where they fail is the templated product. The failure to watch for is a report that arrives weeks after signing with your company name on the header and nothing else specific to you in it. If your tool scores kitchen roles on a different rubric from front desk, the methodology section should say so. Ask to see the contents page of a redacted prior report. If its structure is identical to what's proposed for you before anyone has seen your data, you know what you're buying.
Generalist consultancies and independent statisticians. Two different animals under one heading, joined only by the fact that neither sells bias audits as a product. A large risk practice brings a review chain and an insurance policy behind the signature. An independent statistician or academic brings technical depth with no sales motion attached, at a fraction of the fee. Both are right when the scope is written down tightly and the independence is genuinely clean. The consultancy fails on attention. A mid-size engagement gets partners at the pitch and a second-year analyst for everything after it, and the analyst has never read the statute. The academic fails on a different axis. With an academic the analysis can be excellent and the publication question can still cost you more time than the audit did, because a department that wants a paper out of your candidate data will raise it late rather than early. Ask both types the same question up front: who does the actual analysis, and what happens to the dataset when the engagement closes?
In-house and law-firm coordination. Your own analytics team running the tests, or an employment law firm that owns the legal shape and subcontracts the technical work. Both keep control close. Your data doesn't leave, and the timeline is yours to set. In-house is right when you have senior statistical talent that has run selection-parity work before and the tool's outputs already sit somewhere you can query. The law-firm route is right when you expect the file to be read adversarially, because the firm builds it with that reading in mind. In-house fails on independence before it fails on skill. The people who helped configure the tool cannot audit it, and in a small analytics function that's often the same two people. The law-firm route fails more quietly. Ask which technical auditor the firm intends to bring, and whether your vendor has worked with them before. The pool of people who can do this work is not large. Nobody is being dishonest. Nobody is checking, either. Ask, in writing, who suggested the auditor.
Five Diagnostic Questions
Do you know which of your tools count as Automated Employment Decision Tools under the statute? Don't answer from the software inventory. It lists what finance paid for, not what touches a candidate. Book an hour with a recruiter and walk one live requisition end to end, noting every point where something scores a person or rejects one without a human reading the file. Then ask procurement for anything the recruiter didn't mention. Sourcing tools often arrive on a marketing budget. Guess at the list and the audit becomes your smaller problem: you can't send a notice for a tool you haven't found.
Has the auditor been involved with this tool at all? The rule is about involvement in using, developing or distributing it, with no time limit attached, so ask the auditor in writing, then ask the vendor separately and compare the answers. Check the vendor's partner pages. Search the auditor's name next to the vendor's on the open web. Ask directly about referral fees, the relationship least likely to be volunteered. If any of it turns up something the auditor didn't disclose, you don't have an independent audit. You have a marketing artefact with a signature on it.
Does the audit cover your applicant pool, or the vendor's entire customer base? Find the population count before you read anything else. If it's a round number in the hundreds of thousands, it isn't your pool. Then look for your role profiles by name. A pooled audit tells you something real about the tool in general. It tells you nothing about the nursing candidates you actually screen in Queens. Ask the auditor which of your requisitions were included. If the answer is a category rather than a list, that's your answer.
Can you show where and when a candidate actually sees the notice? The timing and wording sit in the city rules and are worth reading with counsel rather than from memory, but you can test placement yourself rather than asking about it. Apply to one of your own live NYC roles from a personal address on a phone, and note what you saw and when, relative to the assessment invitation. If the notice sits in a careers-page footer, or on a page candidates only reach after they've submitted, the honest answer is no. Then find out who owns that text. It usually lives in a template nobody has edited since the day it was written.
Is the published summary on your website, and is it the current one? Open a private browser window and find it the way a candidate would, starting from your homepage. Time yourself. Then check which audit the file describes and when it was produced. A stale summary sitting live is worse than a gap. It's a specific published statement about a tool you no longer run that way, with a date on it. "We are working on it" isn't a published summary. Name the one person who owns that page.
Who Can Actually Run the Audit
A specialist AI audit firm
A firm whose whole business is bias auditing for employment AI. They know the test statistics and the format the city expects to see, so you aren't funding someone's education in the statute. It earns its place because the output is built for a hostile reader: the scope names your role profiles, and the small-count problems get declared rather than smoothed over. Right when your applicant pool is large and your roles vary enough that a single pooled test would hide something worth knowing. Where it falls short is standardisation. The specialists that scale do it by templating, and a template quietly assumes your tool is configured the way most of their clients configure theirs. If your tool scores two role families on different rubrics and the proposal never mentions it, you're buying an average of your company, not a picture of it. Push for tailored cut points before you sign, not when the draft lands.
A large consultancy risk practice
A big advisory risk team, usually with a statistician subcontracted underneath. What you're buying is process and standing. There's a review chain above the person doing the work, and there's professional indemnity behind the signature. That matters when the summary is going to a regulator or into a board pack, because the name carries weight with readers who can't assess the statistics themselves. Right when the engagement is large enough to hold a partner's attention through delivery. Where it falls short is the gap between the pitch team and the delivery team. The gap to watch for is between the people in the pitch and the person who does the work. Partners sell it. A junior analyst may run it. The independence on the page is the partner's name, but the judgment in the file is whoever actually sat with the data. Ask who does the analysis, and ask before the fee is agreed.
An independent statistician or academic
A quant with no financial relationship to the vendor, working alone or with one collaborator. Frequently the most technically careful option and usually the cheapest, because you're paying for analysis instead of for a brand and three layers of review. It earns its place because this person will argue with you about whether your counts support the finding at all, which is the conversation most engagements skip to protect the timeline. Right when you have a clean dataset and someone internal who can project-manage. Where it falls short is everything around the analysis. There's no house template for the published summary, and no cover if the analyst is ill in week three. Academics carry a second complication. The institution may want to publish, and your candidate data becomes the price of the paper. Settle publication rights and data deletion in the engagement letter, not in month two.
The tool vendor's own published audit
A PDF the vendor sends you, usually within a day, usually at no cost. It earns a place on this list for exactly one reason. It tells you what the vendor believes about its own model, which makes it worth reading before you commission your own work, because it shows you the test statistics the vendor is comfortable with and the ones it steps around. Read it as evidence about the vendor. Don't file it as your compliance record. It falls short on every axis that counts. The auditor has almost certainly done prior work for the vendor, which goes straight at the independence requirement. The population is pooled across the vendor's customers, so it describes a candidate universe that isn't yours. The document is written to protect a sales motion. The vendor's audit isn't your audit. It's the vendor's.
An internal analytics team
Your own data scientists, working from your ATS export. Maximum control of the timeline, and no candidate data leaves your environment. It earns its place because your team knows things an outsider spends three weeks discovering: which requisitions were reposted, which role codes were retired halfway through the year. That knowledge is the difference between a defensible denominator and an invented one. Right when you have senior statistical talent that has run selection-parity work before. It falls short on independence first. If the team helped tune the tool, or sits under the same leader as the recruiters whose numbers are being examined, an outside reader will say so before they've looked at your method. It also falls short under deadline pressure, because internal work competes with whatever else the quarter demands and rarely wins. There's a version that works. Internal team does the exploratory analysis, external auditor does the audit of record.
An employment law firm coordinating a subcontracted auditor
A practice that owns the legal shape of the work and brings in a technical auditor underneath it. The firm writes the scope with an eye on how the file reads if it's ever produced, and it holds the correspondence if the city comes asking questions. It earns its place when you already sense this will be read adversarially, because a complaint has landed or a candidate has written to your general counsel. Privilege over the exploratory work is a genuine advantage, though what you publish is public by design. Your counsel will tell you where that line falls. Where it falls short is the choice of auditor. Firms have a short list of technical people they call, and that short list overlaps with the vendor's. Nobody is being dishonest. Nobody is checking, either. Ask who suggested the auditor, and ask before the engagement letter is signed.
The Decision Table
| Situation | Scale | Setup | Primary Pain | Recommended Starting Point |
|---|---|---|---|---|
| Single tool, single NYC role profile, pooled applicant volume | 200+ applicants per role | Vendor-hosted tool, careers page is the only candidate touchpoint | Notice and publication, not the audit itself | Employment law firm coordinating a subcontracted auditor |
| Multiple roles, varied applicant volumes, in-house ATS | 1,000+ applicants per year | Mix of tools, internal analytics team, no current audit | Statistical scope across roles, not vendor relationships | Specialist AI audit firm with a per-role scope |
| Recent vendor switch, last audit was for the old tool | Any scale | New tool live less than six months | Establishing the new baseline and the candidate notice flow | Large consultancy risk practice with industry references |
| Small applicant pool for a niche clinical role | Under 100 per role | Single tool, one role profile, low candidate volume | Audit may be underpowered at standard cut points | Independent statistician with explicit scope on power |
| National footprint, NYC is one of many sites | 5,000+ applicants per year | Centralised talent ops, standardised notice process | Keeping the NYC notice and publication ahead of other jurisdictions | Specialist AI audit firm plus internal compliance owner |
| Vendor will not share candidate-level data | Any scale | Closed tool, opaque scoring | The audit cannot happen without access | Escalate to legal counsel before signing any renewal |
| Tool is used only after a human screen | Under 1,000 per year | Tool ranks, humans decide | Whether the tool is in scope at all | Local counsel review of the actual decision flow |
| Internal mobility tool for existing employees | Under 500 per year | Tool scores internal candidates for promotions | Whether the statute covers internal use | Local counsel review of the statute's employee language |
The Cost of Getting This Wrong
The penalty is the part everyone quotes. Up to 500 dollars for a first violation, then between 500 and 1,500 dollars for each subsequent one, enforced by the NYC Department of Consumer and Worker Protection. For most companies that number is rounding error. It's also not the cost that should keep you up.
Start with the hiring you don't do. When the publication lapses, the safe move is to switch the tool off, and the switch always happens mid-search. Two hundred applications a week revert to a recruiter reading CVs by hand. Time to fill stretches. Hiring managers promised a better process watch it get worse, and the room remembers that the next time talent ops proposes a tool.
Then the credibility cost, which lands on a single person. Someone said the vendor handled it. That sentence gets repeated in a board pack, where it hardens. Corrections never travel as far as the claim did, and the person who made it spends a year being asked whether anything else was assumed rather than checked.
There's the candidate, too. The one who asked for the audit and got a vendor PDF, then posted the exchange. Nobody in recruiting can answer the replies, so your recruiters spend a quarter opening calls by defending the process instead of selling the job.
Then there's what an audit creates. Numbers, in a file, with a date on them. That's the point, and it's also permanent. A first audit scoped in a hurry becomes the baseline the second gets compared against, so a denominator you'd have fixed in week one is now a trend line you have to explain. You only get one first time.
And the cost that compounds most quietly is treating this as a one-city errand. Illinois HB 3773 took effect on 1 January 2026. It requires notice when AI is used in employment decisions and bars using ZIP codes as a proxy for a protected class. California Civil Rights Council regulations on automated decision systems took effect on 1 October 2025. Take local advice on each. But a team that ran the NYC work as a bespoke scramble rebuilds from nothing every time, while a team that built a repeatable record answers the next one in a fortnight.
So the question isn't whether you can afford the audit. The question is whether you can afford the audit, the publication, the notice, and the record that ties all three together, in case someone asks to see them next quarter.
When You Are Ready to Go Further
If you've read this far, you've done the part most teams skip. You have separated the vendor's job from yours, and you know the audit, the summary, and the notice are three different deliverables that fail on three different schedules.
The next decision is which auditor fits your applicant pool, your role mix, and your timeline. That's a comparison, not a recommendation, and it's the kind of work HROpsLab does for a living. We are a review publication. We don't sell software, and we don't sell consulting. We test the claim against the statute and the city format, and we tell you which options hold up under both.
If you want a starting point, our comparison work on independent AI bias auditors for NYC Local Law 144 is the natural next read. It maps the firms above to the role profiles and applicant volumes where they actually deliver, and it flags the ones whose independence doesn't survive a second look.
Frequently Asked Questions
Does the vendor's annual bias audit count as my audit?
No. The obligation sits with the employer using the tool, not with the vendor that built it. A vendor audit of its own model across its entire customer base isn't an audit of how the tool performs on your applicant pool, and it doesn't satisfy the publication or notice duties. Read it anyway, because it tells you which test statistics the vendor is comfortable being measured on. Then commission your own, scoped to your role profiles. The practical test takes a minute: open the PDF and look for the population it describes. If your own requisitions aren't in it, it isn't describing you.
What does "independent auditor" mean under the law?
It means the auditor has not been involved in using, developing, or distributing the tool being audited. Prior work for you on unrelated systems is generally fine. Prior work for the vendor on this tool isn't, and it's the single conflict that gets challenged most often. Ask for the relationships in writing before you engage, and ask about referral arrangements as well as project history, because money moves between firms that have never worked on the same file. Ask the vendor separately. If the answers don't line up, you've learned more than either answer told you.
What happens if the audit finds bias in my hiring tool?
You publish the result. The audit is a measurement, not a clearance. The statute expects the summary to be public regardless of the outcome, and a favourable internal review that never makes it to the website is the same failure as an unfavourable one that gets buried. A finding is useful: it points at a cut point worth revisiting or a role profile the tool shouldn't touch. The behaviour that causes lasting damage is the quiet re-run, where someone asks for different parameters until the numbers improve. Those requests live in email, with timestamps. Take local advice on how you describe what you found.
Does Local Law 144 apply to remote roles for a New York City position?
If the role is based in New York City, the law applies to the screening of candidates for that role, regardless of where the candidate is sitting when they complete the assessment. The anchor is the position, not the person. That catches a case teams routinely miss: the fully remote role carrying an NYC office address, often the largest applicant pool in the company. If you scoped your notice by where applicants live rather than where the job sits, check it this week. Confirm the specific facts of your case with local counsel before relying on this in a policy.
How much does a compliant bias audit cost?
It depends on the size of your applicant pool and the number of role profiles you want covered. Get at least two written quotes that name the test statistics and the publication format, then compare the scope rather than the fee. One quote may price a single pooled test across everything you run. Another may price each role profile separately, with an explicit check on whether your counts support a finding. Those are very different products, and they often carry similar headline numbers. Don't budget from a vendor blog post.
What exactly do I have to publish, and where?
A summary of the bias audit results, on your website, in a location that a candidate could reasonably find. Test that last part instead of assuming it. Open a private browser window and see how long it takes to reach the summary from your homepage. A link below a cookie banner on a page candidates only reach after applying is not somewhere anyone finds. The format the city expects to see includes specific elements, so check the latest DCWP guidance and have local counsel sign off before you publish. Then put one named person in charge of keeping that page current.
Is there a small-employer exemption?
The statute doesn't carve out a small-employer exemption by headcount. The trigger is whether you use an Automated Employment Decision Tool to screen or rank candidates or employees for NYC-based positions. Headcount is not what the trigger turns on, so a small employer should check its own position with counsel rather than assuming size settles it. Small employers do face a real practical problem, though it isn't an exemption: with a modest applicant pool, an audit at standard cut points can come back underpowered, and an honest report says so. That's a conversation to have with the auditor at the start, not a discovery to make when the draft arrives.
HROpsLab is an independent review publication for HR operations teams. We test vendor claims, we don't sell software, and we don't sell consulting.