Short answer

Public AI recruiting case studies do not support a single verdict that automation works or fails. They document different systems at different stages: an internal experiment, a redesigned graduate funnel, a discriminatory rule that produced automatic rejections, an active lawsuit over screening software and an internal talent platform described by its developer.

The cases are useful only when their evidence types remain separate. A regulator’s consent decree is stronger evidence of a resolved enforcement matter than a vendor case study is of customer return on investment. A court order can establish what claims may proceed without proving those claims. A company careers page can document a current process but not historical model performance.

This revision was checked on September 13, 2026. It removes the previous article’s unsupported $2.3 billion aggregate, fictional applicants, invented executive interviews and precise savings presented without a traceable source. Digidai did not interview any company, candidate, vendor or regulator discussed here.

Amazon: a reported experiment, not a production benchmark

Reuters reported in 2018 that Amazon had developed an experimental resume-ranking system using patterns in historical applications. According to the report, the system learned signals that disadvantaged resumes containing terms associated with women, and the project was abandoned after attempts to correct it did not give the team confidence that other discriminatory patterns were gone.

The often-repeated summary that “Amazon used AI to reject women” loses important scope. Reuters described an internal project and said Amazon recruiters looked at recommendations generated by the tool, but the company said it was never used to evaluate candidates in the active hiring process. Amazon did not publish a model card, training-data description or validation report that would let an outside reviewer reproduce the system.

The case still supplies a durable lesson. Historical outcomes are not neutral labels. A model trained to resemble prior hiring can encode occupational segregation, past sourcing choices and the language of previous applicants. Removing an explicit term such as “women’s” does not prove that correlated proxies have disappeared.

The correct response is not a promise to remove protected attributes. It is a documented review of features, labels, subgroup outcomes, job relevance and failure cases, followed by a decision to change or stop the system when the evidence is inadequate. Amazon’s reported choice to abandon the experiment is itself an important control outcome.

Unilever: an old transformation claim and a current process

A LinkedIn Talent Solutions 2018 recruiting trends report presented Unilever’s early-career hiring redesign as a case study. It described a mobile application, game-based assessment, recorded video interview and in-person Discovery Centre. The report said the company cut hiring time by 75%, expanded the US school pool and moved candidates through the process faster.

Those figures are useful as a historical company case, but the document is not an independent experiment. It does not publish a comparison group, complete selection-rate tables, attrition by subgroup or a long-term quality-of-hire method. “Time cut” also needs defined start and end points before another employer can compare it with its own process.

The current workflow has changed by region. Unilever’s 2026 UK and Ireland Future Leaders Programme page describes profile assessments, a digital interview and a Discovery Centre. It explicitly says AI is not used as a screening tool in that program and a person reviews the digital interview. Other current Unilever country pages describe different sequences.

That update prevents a common error: treating a 2017 deployment story as a permanent, global description of Unilever hiring. A buyer or candidate needs the country, program, year, vendor and decision role. A recorded interview is not necessarily scored by AI, and an online assessment is not necessarily generative AI.

The useful measure is the complete funnel. Faster screening can improve candidate responsiveness, but it can also move a bottleneck into assessment completion, manager review or scheduling. Track qualified applicants entering each stage, time in stage, withdrawals, accommodations, offers and accepted starts.

iTutorGroup: simple rules can cause automated harm

The Equal Employment Opportunity Commission’s 2023 settlement announcement says iTutorGroup companies programmed application software to reject female applicants aged 55 or older and male applicants aged 60 or older. The agency said more than 200 qualified US applicants were rejected. A consent decree provided $365,000 and additional relief.

This is an enforcement settlement, not a finding after trial. The public record nevertheless supports the core facts of the resolved case and the obligations in the decree. It does not require speculation about an individual applicant’s life or a private conversation with a company engineer.

The system also demonstrates why “AI bias” is sometimes an imprecise label. A deterministic age threshold can discriminate without a complex model. Governance must cover business rules, knockout questions, filters and integrations as well as machine learning. A review limited to model training data can miss the rule that actually sends a rejection.

Test boundary values directly. Submit records immediately below, at and above each threshold. Vary one field at a time. Confirm which rule fired, what the candidate received, whether a person can review the outcome and how the employer disables the rule. These tests are inexpensive compared with reconstructing a hidden filter after complaints arrive.

Workday: litigation and a vendor audit answer different questions

Mobley v. Workday is an ongoing federal case in which plaintiffs allege discrimination based on race, age and disability in algorithmic applicant screening. A July 2026 court order addressed which allegations remained in an amended complaint. Those procedural rulings are not a finding that the plaintiffs proved discriminatory impact.

The complaint remains disputed. Separately, Workday says its products support human decisions rather than automatically reject candidates. That position is a company statement. It should be presented alongside, not substituted for, the procedural record.

In 2026 Workday published an analysis of its own HiredScore Spotlight deployment. Workday says an independent firm analyzed applicants for five high-volume job profiles in the greater New York City area during a six-month period and found no evidence of disparate impact in the reported ratios. The page also says the result is specific to Workday’s implementation and is not intended to satisfy a customer’s legal obligation.

That limitation is exactly right for evaluation. A result from Workday’s applicant pool, job profiles, configuration and human workflow cannot certify a different customer’s deployment. The customer still needs its own population data, job analysis, selection process and production monitoring. Conversely, allegations in a lawsuit do not prove that every use of the product is unlawful.

The case exposes a practical procurement issue: evidence access. An employer should contract for model and feature descriptions, relevant test summaries, version notices, records needed for its own audit and cooperation when a candidate challenges a decision. Responsibility cannot be managed if the deployer and developer each holds half the record and neither can reconstruct the outcome.

IBM: internal talent AI is a company-reported case

IBM describes Wf360 as an internal HR data and AI platform. Its company case study says the platform combines more than 30 sources, provides job recommendations to 180,000 employees and moved monthly insights to HR leaders 23 days earlier. It says flexible APIs made delivery of new AI solutions seven times faster.

These are IBM’s own operational claims about a system IBM built and uses. The page is useful because it names the data-quality problem, system scope and some workflow boundaries. It is not a controlled study and does not publish enough underlying data to verify the productivity multiplier independently.

Internal mobility also differs from external hiring. The organization may have richer skills, job and performance records for employees, but using them raises purpose, access and fairness questions. A recommendation can widen visibility into roles or reinforce a manager’s prior labels. Employees need to know which records feed the match and how to correct inaccurate skills.

IBM says managers make final salary decisions when another Wf360 feature offers compensation recommendations. That boundary should be tested in practice. Record the recommendation, evidence shown, manager action and override. A final human click is weak assurance if managers cannot inspect or challenge the basis.

What survives across the five records

The cases do not share one definition of AI, one outcome or one evidence standard. Their common lesson is that the decision pipeline matters more than the product label.

Public recordWhat it supportsWhat it does not establish
Reuters account of AmazonA reported internal experiment and its reported failure modePerformance of a published or current Amazon hiring product
LinkedIn and Unilever pagesA historical redesign claim and a current regional processA permanent global workflow or independent ROI estimate
EEOC iTutorGroup decreeTerms and allegations resolved through a public settlementA trial verdict about a machine-learning model
Workday court recordProcedural status and the allegations before the courtLiability or performance across all customer deployments
Workday Spotlight analysisReported ratios for a defined Workday deploymentCompliance or fairness for another employer
IBM Wf360 case studyIBM’s description of internal data and talent workflowsIndependently verified productivity or hiring quality

Academic research adds another caution. The NBER paper “Systemic Discrimination: Theory and Measurement” shows how disparities can accumulate across multiple stages and how language in recommendation letters can carry gender-related differences into later decisions. It does not evaluate the five company systems above. It supports testing the pipeline rather than declaring fairness from one stage.

Replace promotional metrics with an evidence stack

Speed and labor savings matter, but they should sit below eligibility, validity and completed hiring outcomes. A practical measurement stack has four levels:

  1. System integrity: eligible candidates enter the workflow once, stages and dispositions are recorded, integrations reconcile and errors are visible.
  2. Candidate access: assessments work with supported assistive technology, accommodation requests reach a person and notices appear before the relevant use.
  3. Decision quality: criteria map to the job, reviewers receive useful evidence, subgroup differences are investigated and overrides are sampled.
  4. Business outcome: qualified hires start, remain through a defined period and meet a predeclared performance measure without unacceptable correction or complaint costs.

Do not combine the levels into one “AI ROI” number. A tool can reduce recruiter minutes while losing qualified applicants. A broader applicant pool can look diverse while later stages reproduce the old outcome. A high completion rate can coexist with an inaccessible path for a smaller group.

Use a frozen baseline and a parallel or staged rollout where practical. Define the job population, period and success measure before seeing the result. Preserve candidate counts at every stage, not percentages alone. Document other changes such as advertising, compensation or recruiter staffing that could explain an apparent improvement.

An agent-friendly hiring system needs receipts

Recruiting agents make the evidence problem more urgent because they can retrieve data, recommend a disposition and take actions across systems. A safe design separates planning from execution.

The agent should use stable requisition and candidate identifiers, retrieve the current approved job criteria, state which evidence supports a recommendation and expose uncertainty. It should not infer protected or sensitive traits to fill gaps. Write actions should be restricted by role and require approval for consequential changes.

Every action needs a receipt: model and prompt version, inputs used, tools called, reviewer, final disposition, messages sent and downstream system response. A statement that the agent “scheduled the interview” is incomplete until the calendar event exists for the right people and the candidate received the correct notice.

Retain failure examples as well as success samples. Test duplicate applications, missing credentials, interrupted assessments, changed availability, accommodation requests and documents in supported languages. A production gate should fail when the system silently drops a candidate, cannot explain the controlling criterion or cannot reverse an incorrect write.

Correction and source scope

The September 13, 2026 revision withdraws fictional candidate stories, purported first-person interviews, anonymous practitioner quotations, invented internal conversations and unsupported company savings. It removes the $2.3 billion headline because the prior article did not provide a defensible calculation or source.

It also narrows several common claims. Amazon is described as a Reuters-reported experiment rather than a deployed rejection system. Unilever’s older transformation metrics are labeled as a case-study claim and compared with a current regional careers page. iTutorGroup is described through the EEOC settlement. Workday allegations remain distinct from court findings and Workday’s own audit disclosure. IBM figures remain company claims.

The original file name, publication date and URL remain unchanged. The evaluation framework is Digidai analysis and is not attributed to the companies, candidates, court or regulator.