The Interview Becomes a Model Interface
AI hiring tools do not merely screen applicants. They change what a job seeker must become legible to before a human institution will answer.
A hiring model interface is the institutional boundary through which applicant-supplied or derived evidence enters automated processing and the resulting parse, score, rank, summary, or flag changes what happens next. It may be visible, like a chatbot or recorded assessment, or hidden, like a parser that controls a review queue.
The source record, model, selection rule, and workflow are separate layers. Governance fails when they are collapsed into the single phrase “the AI,” because an accurate model cannot repair a false record, validate a cutoff, make an inaccessible test fair, or create an effective appeal.
The Gate Before the Gate
The conventional hiring story treats the interview as the threshold: a person sits across from other people, performs competence, answers questions, and tries to become legible as a worker. In practice, gatekeeping has always started earlier through networks, credentials, application forms, and clerical screens. Those practices carried class codes, race codes, gender codes, disability barriers, accent penalties, and ordinary human bias.
Model-mediated hiring scales and obscures that earlier gate. A resume parser, ranked shortlist, chatbot prescreen, recorded-interview analyzer, skills assessment, targeting system, or automated rejection rule can determine who receives human attention. The applicant may never learn which data, score, model, vendor, threshold, or reviewer blocked the path.
This is not a small administrative change. Work is one of the main ways a society distributes money, status, health insurance, schedule control, immigration stability, housing access, apprenticeship, and adult identity. A system that shapes who gets seen by an employer is not just a productivity tool. It is a civic gate.
The legal layer now names parts of this interface. New York City's Automated Employment Decision Tools law imposes audit, publication, and notice duties on a defined class of tools. Illinois regulates AI analysis of applicant video and separately prohibits discriminatory employment uses of AI. Federal civil-rights and disability law continues to govern selection procedures. The EU AI Act classifies many employment and worker-management systems as high risk, while prohibiting a narrower class of workplace emotion-inference systems.
The pattern is clear: the hiring interface has become a governance problem. The ethical and operational boundary is broader than any one statute, however; a tool can materially control attention without meeting a particular jurisdiction's definition of covered AI.
What Is Being Automated
"AI hiring tool" is too broad a phrase. It can mean systems that place job ads, parse resumes, infer skills, rank candidates, score tests, conduct chatbot prescreens, analyze recorded interviews, recommend questions, summarize notes, predict attrition, infer “fit,” or recommend promotion and termination decisions.
Four objects should be kept separate. The source record contains applicant statements, imported records, assessments, and derived fields. A model transforms those inputs into a score, classification, recommendation, or summary. A selection procedure decides how that output matters: a cutoff, rank order, knockout rule, reviewer queue, or comparison. The workflow determines what the applicant sees, how accommodation works, who reviews the evidence, what is retained, and whether an error can be corrected. A model can be accurate on its test set while the record is wrong or the cutoff is invalid; a measured procedure can still sit inside an inaccessible or unappealable workflow.
Each layer has a different failure mode. Targeting can shape who learns about an opportunity. Parsing can omit a qualification or convert an inference into an apparent fact. A similarity model can compare a worker's history with an unexamined ideal. A timed test can measure an impairment instead of the skill it claims to assess. Video or voice analysis can turn disability, accent, lighting, bandwidth, speech pattern, or camera comfort into a proxy for merit. A generated summary can invent or overstate a trait. A rank can frame later human judgment even when it never issues the final rejection.
Material influence is therefore wider than automatic rejection. A tool shapes a decision when it determines who receives an invitation, which file appears first, what evidence a reviewer sees, what question is asked, whether an exception is escalated, or how much scrutiny a candidate receives. The employer should map those state transitions rather than inventory only tools that produce a final yes or no.
Upturn's 2021 report on large hourly employers found that hiring technologies had become part of the application process for low-wage work, including online assessments and applicant-tracking systems. That matters because automated hiring is not only an elite white-collar problem. It sits inside retail, logistics, food service, warehousing, customer support, care work, and the everyday search for stable hours.
The practical asymmetry is severe. Employers and vendors can inspect the funnel. Applicants see a form. The institution can tune a threshold, change a prompt, update a model, or replace a vendor. The rejected person may receive no useful reason, no record of which layer mattered, no correction route, and no evidence that the procedure was connected to the job.
That is why hiring automation is not only about bias in a technical sense. It is about legibility under unequal power. The applicant must become machine-readable. The employer does not have to become applicant-readable.
Current Context
As of August 12, 2026, AI hiring governance is a jurisdiction-specific patchwork of civil-rights law, selection-procedure rules, city audit duties, state employment and consumer protections, privacy and background-report law, private litigation, and staged European requirements. “Uses AI” is not a legal conclusion. Coverage turns on facts such as location, who deploys the system, what data it uses, whether it materially influences a consequential decision, and whether a person can be identified or affected.
New York City's Local Law 144 covers a defined automated employment decision tool that substantially assists or replaces discretionary decision-making in hiring or promotion. DCWP says a covered employer or employment agency must obtain an independent bias audit no more than one year before use, publish a summary, and provide required notice—normally ten business days before use. Its FAQ excludes scans of a resume bank, outreach, and invitations to apply when the person has not applied for a specific position. The law is therefore narrower than this essay's model-interface boundary and narrower than “all hiring AI.”
The required New York audit is also a bounded artifact, not a general safety certificate. DCWP permits pooled historical data under stated conditions and test data when historical data is insufficient; it does not require pooled employers to have hired for the same type of position. The summary must disclose the data source and explanation, so an employer should test whether the audit population, job, version, configuration, and threshold actually match its use. A 2025 New York State Comptroller audit separately found complaint-routing and enforcement gaps.
State rules reach different objects. California regulations effective October 1, 2025 clarify that automated-decision systems may violate employment-discrimination law, bring automated-decision data within a four-year employment-record rule, and warn that some assessments may elicit prohibited medical information. Illinois's video-interview law requires pre-interview notice, an explanation of the AI analysis, consent, limited sharing, and deletion of recordings on request; Public Act 103-0804, effective January 1, 2026, separately prohibits discriminatory AI use across listed employment functions and directs the state civil-rights department to implement employer notice rules. The department published proposed implementing rules in May 2026; a proposal is not a final rule.
Colorado's enacted SB26-189 applies from January 1, 2027 to covered automated decision-making technologies that materially influence consequential decisions, including employment. Among other duties and exceptions, it requires developer documentation, deployer notice, a plain-language description after an adverse outcome, access to and correction of factual personal data, meaningful human review and reconsideration on request, and at least three years of compliance records. It is enforced by the attorney general and does not create a new private right of action.
In the European Union, Annex III of the AI Act lists many recruitment, candidate-evaluation, worker-management, task-allocation, monitoring, and performance-evaluation uses as high risk. Regulation (EU) 2026/1744, in force since July 27, 2026, moved the relevant high-risk requirements from August 2026 to December 2, 2027. When those duties apply, Article 26 will require deployers to assign oversight to people with competence, training, authority, and support; monitor the system; retain controlled logs; suspend risky use; and inform affected people. The separate Article 5 prohibition on workplace emotion inference, except for medical or safety reasons, has applied since February 2, 2025.
The current context therefore is not “AI hiring is illegal” or “AI hiring is solved by audits.” System inventories, audit summaries, notices, validation records, logs, and applicant data have become compliance evidence. A serious employer must know which component shaped the decision, which rule applies, whether the selection procedure is job-related, whether accommodation and correction worked, and whether the record can support later review.
The Civil-Rights Layer
The United States does not need a new civil-rights theory to recognize a model-mediated screen as a selection procedure. Under Title VII disparate-impact analysis, a neutral procedure that disproportionately excludes people by race, color, religion, sex, or national origin can require evidence that it is job-related and consistent with business necessity; a challenger may still identify a less discriminatory alternative. The Uniform Guidelines describe validation disciplines. An impact ratio can flag a difference, but it does not by itself establish validity or legal compliance.
Reliability, validity, and fairness are not synonyms. A score can be repeatable without measuring a job-relevant construct. A procedure can correlate with an outcome on average while failing for the role, population, language, accessibility setting, or threshold in which it is used. A favorable group ratio cannot establish that the construct is lawful or useful, and an unfavorable ratio does not reveal which layer caused the disparity. Job analysis, criterion definition, validation evidence, subgroup testing, and workflow review answer different questions.
The Americans with Disabilities Act adds distinct duties. A tool can screen out a qualified person because it measures an impairment instead of the job factor it purports to measure, lacks an effective accommodation process, or elicits disability-related or medical information before a conditional offer. The DOJ's applicant guidance gives concrete examples involving game-based tests and facial or voice analysis. Accessibility must be designed before the timed assessment or recorded interview begins; a help address that answers after rejection is not an accommodation path.
A separate federal layer can attach when an employer obtains an employment background report from a third party. If the product is a “consumer report” under the Fair Credit Reporting Act, the employer must follow disclosure and authorization rules and, before adverse action based on the report, give the person a copy and a summary of rights; later notice must identify the reporting company and dispute rights. Not every in-house score is a consumer report. The point is to classify the data product instead of assuming that an AI label displaces older accuracy and correction law.
These regimes ask related but different questions: Was the procedure discriminatory? Was it validated for this job? Did it measure disability rather than ability? Was the applicant told about a covered report and able to correct it? Was the interface accessible? A single vendor “fairness” badge cannot answer all five.
Hiring systems are high stakes not because every model is malicious, but because employment decisions shape livelihood, information is asymmetric, and an error can remain invisible to the person excluded.
Bias-Audit Theater
New York City's rule is important because it makes an independent audit summary and advance notice public-facing conditions of use for a covered tool. Its minimum audit calculates selection or scoring rates and impact ratios across sex, race or ethnicity, and intersectional categories. The summary's data source, exclusions, sample sizes, unknown categories, job mix, version, and audit date determine what those ratios can support. DCWP's FAQ is equally important about the limit: the law does not prescribe a particular response when an audit indicates disparate impact. Other anti-discrimination law still controls what the employer must do.
The weakness is visible in enforcement. The 2025 New York State Comptroller audit found an ineffective complaint-routing process, no recent public outreach, limited use of technical support, and a gap between DCWP's review of 32 companies and the Comptroller's identification of at least 17 potential instances of non-compliance. Audit law needs enforcement capacity and a usable complaint route, especially when a person cannot report a tool they were never told existed.
An audit can also become ritual if it measures group selection rates without testing the construct or the workflow. It may not show whether the procedure predicts relevant performance, whether a threshold is defensible, whether small samples hide intersectional harms, whether disabled applicants can complete the process, whether a model change invalidated prior results, whether reviewers over-rely on scores, or whether rejected applicants can correct an error.
Research results are system- and method-specific. Wilson and Caliskan's 2024 simulated retrieval study tested three embedding models across nine occupations and found race, gender, and intersectional disparities, with particularly adverse patterns for Black men. A June 2026 preprint by Gao, Jiang, and Yan sent paired synthetic profiles to fourteen generative language models and found a pro-White gap only in the oldest model they tested; later models showed either no measured gap or a reversal. That experiment produced model “yes” or “no” outputs, not callbacks from employers. Neither result certifies a deployed product. Together they show why an employer must test the actual model version, prompt, documents, threshold, role, and population—and why the direction of disparity is not the same thing as job validity.
“Human in the loop” is not an audit result. In a controlled 2025 resume-screening experiment with simulated race-favoring recommendations, Wilson and colleagues found that the recommendations shifted participant choices even when some participants rated them poorly. Their 2026 analysis of the same experiment found that participants spent up to 55.6 percent longer viewing resumes when no recommendation was shown. This is not an independent field replication or an estimate of recruiter behavior, but it demonstrates that oversight needs time, source access, authority to disagree, and monitoring of review behavior, overrides, and outcomes.
Responsibility also crosses a chain: a model provider supplies a component; a hiring vendor configures it; an employer chooses criteria and cutoffs; a recruiter acts on its output. Governance must assign each party documentation, testing, incident, change-notice, and correction duties. Otherwise the audit becomes a polished PDF floating above the actual decision. This is a procurement problem as much as a model problem, as the pages on AI procurement and vendor and platform governance explain.
The ongoing Mobley v. Workday litigation shows the legal pressure point. Applicants allege race, age, and disability discrimination by algorithmic screening tools; Workday disputes the claims. The district court held at the pleading stage that allegations could support an employer-agent theory, and in May 2025 preliminarily certified an ADEA collective for notice purposes. Later rulings have allowed parts of amended claims to continue. None is a merits finding that the tools discriminated. The structural question is whether responsibility can follow a vendor when its system allegedly does more than transmit an employer's fixed rule.
The Interface Discipline
The most under-discussed part of model-mediated hiring is behavioral. Applicants who know or suspect that a parser, ranker, or generated score stands between them and a recruiter may adapt to an imagined machine.
Some rewrite resumes around keywords, mimic job descriptions, buy optimization services, rehearse recorded answers under unfamiliar constraints, or use generative tools to standardize their language. These tactics are not evidence of deception by themselves. They are predictable responses to an interface whose criteria are consequential but unclear.
The interface therefore changes the thing it claims merely to measure. It teaches applicants which formats, histories, gaps, credentials, words, and performances appear legible. If employers then train or tune on those adapted applications, apparent “fit” can become evidence of interface fluency rather than ability to do the work.
Conversational polish is also weak evidence of interview quality. A separate field study of bot-led qualitative research interviews—not employment screening—found that topic coverage and fluent exchange did not guarantee deep follow-up. The domain difference matters, but so does the narrow lesson: a hiring chatbot should be evaluated for what it elicits, misses, summarizes, and allows a person to correct, not for whether it sounds attentive.
Interface failures must remain system events rather than negative candidate evidence. A dropped connection, speech-recognition error, timeout, camera failure, repeated answer, request for clarification, or accommodation handoff should not silently lower a score or be interpreted as motivation, integrity, or competence. The record should show the interruption, its technical source where known, what evidence was lost, and whether the candidate received a no-penalty restart or equivalent alternative.
If models help write applications while other models filter them, the labor market can enter a recursive loop: generated applications are judged by generated representations, and both sides optimize the surface. The answer is not to guess which applicant used AI. It is to make role criteria observable, test actual work where appropriate, and keep a verified human route for unusual but qualified candidates.
The Candidate Record
The practical governance object is a candidate decision record: bounded decision provenance, not a permanent dossier. It lets an applicant, employer, auditor, regulator, or court reconstruct what evidence entered, which transformations occurred, what rule changed the candidate's state, and who could reverse it—without demanding source code or irrelevant deliberation.
A usable record should separate the role criteria from the tools used; the applicant's own words from parsed or inferred fields; assessment inputs from model outputs; generated summaries from human notes; the vendor recommendation from the employer's action; and protected demographic audit data, accommodation information, and medical information from routine recruiter access. It should identify the system and vendor, model or rules version, prompt or configuration where material, data sources, score or recommendation, threshold, state transition, reviewer, override, notice, correction request, and final disposition.
This is where AI audit trails and notice and appeal become concrete. A rejected applicant does not need a mystical explanation of model internals. They need to know whether the system used the right record, whether an accessibility barrier distorted the score, whether a generated summary invented a trait, whether a resume parser dropped relevant work, and whether a human could correct the result. An automatically generated reason written after the decision is not provenance unless it is tied to the actual rule and evidence that changed the outcome.
Retention must be purpose-specific. A legal hold or employment-record rule may require preservation, while privacy and deletion rules may require a different treatment of recordings, audit attributes, vendor telemetry, or consumer reports. The institution should document which rule controls each category, restrict access, and delete data when its stated purpose and legal basis end. “Future model improvement” is not a retention schedule.
The record also supports security without converting suspicion into automatic rejection. If untrusted resume text can manipulate an LLM screener, the system should isolate instructions in applicant documents, preserve the original, log the transformation, and route suspected manipulation to review. The resume prompt-injection analysis explains that threat boundary. A security flag should remain distinguishable from an applicant's qualifications and open to correction.
A Better Standard
A serious hiring-governance standard should start from the applicant's position, not the vendor's promise.
First, inventory material influence. Name every component that targets, parses, enriches, tests, scores, ranks, summarizes, routes, or rejects; record its owner, purpose, input source, output, downstream rule, affected roles and people, and legal classification. A “decision-support” label does not remove a system that controls reviewer attention from scope.
Second, give decision-specific notice. Before the relevant interaction, tell applicants what kind of system will be used, which decision it supports, what data categories it uses, how to request accommodation, and where to seek correction or human review. Generic privacy boilerplate is not operational notice.
Third, validate the selection procedure for the role. Start with a documented job analysis and the construct the procedure is supposed to measure. A model's ability to rank does not show that its output measures an essential job function or predicts relevant performance. Validate the model-plus-threshold for the position and intended population, compare less discriminatory alternatives, and reassess after material changes.
Fourth, audit the workflow, not only group ratios. Test intersectional outcomes, sample size, missing demographic data, proxy variables, false negatives, accessibility, accommodation timing, generated-summary accuracy, reviewer reliance, overrides, and who reaches a human. Publish enough method to make a result interpretable.
Fifth, make accommodation usable before the screen. The request route must be accessible, timely, staffed, and separated from evaluation. An alternative should measure the same job factor without converting the request into a negative signal.
Sixth, define meaningful human review. The reviewer needs the underlying evidence, model limitations, time, competence, authority to disagree, and a recorded reason for adopting or reversing a consequential output while the decision is still reversible. Measure review behavior; do not infer oversight from the presence of a human name.
Seventh, provide correction and reconsideration. Distinguish correction of factual or parsing errors, accommodation review, and reconsideration of the selection decision. A person should be able to contest a generated summary and reach someone able to change the outcome. Give an accurate reason, set response times, keep the opportunity meaningful while review occurs, and preserve non-retaliation.
Eighth, keep bounded decision provenance. Preserve what is legally and operationally necessary, apply access controls and retention schedules by data category, and separate applicant statements, inferred data, audit attributes, accommodations, model outputs, and final decisions.
Ninth, prohibit irrelevant inference. Facial expression, voice tone, eye contact, keystroke rhythm, response speed, camera presence, or inferred emotion should not become an employment test without a lawful, validated, role-specific basis. Some uses are prohibited outright in some jurisdictions.
Tenth, make vendor duties contractual and testable. Require documentation, role-specific validation cooperation, accessibility evidence, model and prompt change notices, incident reporting, audit access, data export and deletion, subprocessors, and a remedy when the vendor's evidence is inadequate.
Eleventh, set monitoring and stop conditions. Track drift, outages, unexpected group effects, overrides, complaints, accommodation failures, and false-negative samples. Assign authority to pause the procedure without waiting for a lawsuit or annual audit.
The standard is simple: a hiring system should make the institution more answerable, not only the applicant more measurable.
Source Discipline
Sources about AI hiring answer different questions. Statutes and current regulations establish legal text; agency FAQs and technical assistance state interpretations but may not bind a court; audit reports assess administration; court orders establish procedural rulings, not the truth of allegations; controlled experiments test specified conditions; vendor material describes a product from the seller's position. Claims should not migrate between those categories without qualification.
This essay's operational category is intentionally broader than a covered AEDT, high-risk AI system, consumer report, or automated decision-making technology under any one law. The New York FAQ, for example, excludes certain outreach and resume-bank activity even though those systems can shape who gets the chance to apply. A governance inventory may be broader than legal coverage; it should not imply that every inventoried tool carries every cited legal duty.
Dates matter. New York City began enforcing Local Law 144 on July 5, 2023. California's employment regulations took effect October 1, 2025. Illinois Public Act 103-0804 took effect January 1, 2026, while implementing rules proposed in May 2026 should not be cited as final. Colorado's SB26-189 duties begin January 1, 2027. Regulation (EU) 2026/1744 moved Annex III high-risk requirements to December 2, 2027; the workplace emotion-inference prohibition already applies. Current legal and policy claims here were checked against primary sources on August 12, 2026.
The research does not support one timeless claim that “LLMs are biased” or “new models fixed bias.” The 2024 retrieval study, the 2025 human-AI experiment, its 2026 follow-up analysis, and the 2026 paired-profile preprint use different systems, tasks, populations, and outcomes. Gao, Jiang, and Yan measured model outputs in a controlled simulation, not employer callbacks; Wilson and colleagues' 2026 paper reanalyzed the earlier experiment rather than independently replicating it. Their disagreement is a reason for versioned, deployment-specific evaluation, not a reason to select the most convenient paper. Upturn's report remains useful for its documented view of hourly-work application processes, but it is not an internal audit of every employer or vendor.
The Workday case is cited for the developing question of vendor responsibility. Pleading rulings and preliminary collective certification let claims proceed and notice issue under specified standards; they do not establish that the challenged tools caused unlawful discrimination. Workday disputes the allegations.
Source discipline also means not overreading compliance artifacts. A city bias-audit summary is not a validation study. A four-fifths calculation is not a safe harbor. A vendor explainability memo is not an appeal process. A model card is not evidence that a particular threshold is job-related. A human reviewer is not proof of meaningful review.
What This Changes
The job application is a ritual of recognition. A person asks an institution to see them as useful, trustworthy, trainable, and worth admitting into its economy. AI changes that ritual by placing a model between the person and the institution.
The danger is not only that the model may be biased. The danger is that the model can become the institution's first imagination of the person. It receives the resume before the manager. It ranks the candidate before the conversation. It turns a life history into features, similarities, probabilities, and thresholds. Then the institution treats that representation as if it were neutral contact with reality.
That is recursive reality at the labor gate. The model describes the applicant. The institution acts on the description. Applicants adapt to the description system. Future data records the adapted behavior. The next model learns from the world the previous interface helped produce.
Good governance interrupts that loop. It asks what the system measured, why the measurement matters, who was excluded, what alternatives exist, how a person can appeal, and whether the tool actually improves the human practice it entered.
Work should not require obedience to an unseen classifier as the price of being considered. If the interview becomes a model interface, the interface must be named, tested, limited, and made answerable to the people whose futures it filters.
Related Pages
- AI in Employment
- Opaque Scoring Systems
- Human Oversight of AI Systems
- Algorithmic Impact Assessments
- AI Audits and Third-Party Assurance
- AI Audit Trails
- Notice and Appeal
- Algorithmic Recourse
- AI System Inventory
- AI Post-Market Monitoring
- Automation Bias
- Data Minimization
- AI Procurement
- Vendor and Platform Governance
- Privacy and Data Stewardship
- The Résumé Becomes the Prompt Injection Payload
- The Bot-Led Interview Becomes the Listening Protocol
- The AI Clause Becomes a Workplace Constitution
- The Emotion Detector Becomes a Workplace Polygraph
Sources
- New York City Department of Consumer and Worker Protection, Automated Employment Decision Tools, reviewed August 12, 2026.
- New York City Department of Consumer and Worker Protection, Automated Employment Decision Tools: Frequently Asked Questions, June 29, 2023.
- Office of the New York State Comptroller, Enforcement of Local Law 144 - Automated Employment Decision Tools, December 2, 2025.
- Illinois General Assembly, Artificial Intelligence Video Interview Act, 820 ILCS 42.
- Illinois General Assembly, Public Act 103-0804, effective January 1, 2026.
- Illinois General Assembly Joint Committee on Administrative Rules, The Flinn Report, Volume 50, Issue 20, proposed AI hiring-decision rules, May 15, 2026.
- California Civil Rights Department, Civil Rights Council Secures Approval for Regulations to Protect Against Employment Discrimination Related to Artificial Intelligence, June 30, 2025; final regulatory text.
- Colorado General Assembly, SB26-189: Automated Decision-Making Technology, signed May 14, 2026.
- U.S. Equal Employment Opportunity Commission, Employment Tests and Selection Procedures, December 1, 2007.
- Electronic Code of Federal Regulations, 29 CFR Part 1607: Uniform Guidelines on Employee Selection Procedures, current official text.
- U.S. Department of Justice Civil Rights Division, Algorithms, Artificial Intelligence, and Disability Discrimination in Hiring, May 12, 2022.
- Federal Trade Commission, Using Consumer Reports: What Employers Need to Know, October 2016.
- European Union, Regulation (EU) 2024/1689 (Artificial Intelligence Act), especially Article 5, Article 26, and Annex III.
- European Union, Regulation (EU) 2026/1744, especially Article 1(40), July 24, 2026.
- Upturn, Essential Work: Analyzing the Hiring Technologies of Large Hourly Employers, May 2021.
- Kyra Wilson and Aylin Caliskan, Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval, Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society 7(1), 1578–1590, 2024.
- Kyra Wilson, Mattea Sim, Anna-Maria Gueorguieva, and Aylin Caliskan, No Thoughts Just AI: Biased LLM Hiring Recommendations Alter Human Decision Making and Limit Human Autonomy, Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society 8(3), 2692–2704, 2025.
- Kyra Wilson, Mattea Sim, Anna-Maria Gueorguieva, Soham Chatterjee, and Aylin Caliskan, Resume Screening, Fast and Slow: (Biased) AI Recommendations' Influence on Human Decision Making, ACM FAccT, 2026; open manuscript.
- Zhenyu Gao, Wenxi Jiang, and Yutong Yan, Can LLMs Hire Fairly? Racial Bias in Resume Screening, arXiv:2606.28978v1, submitted June 27, 2026.
- He Zhang, Kambinachi Chukwuma, ChanMin Kim, and John M. Carroll, When the Interviewer Is a Bot: Behavior, Breakdowns, and Trust in MLLM-Led Interviews, arXiv:2608.10412v1, submitted August 11, 2026.
- U.S. Equal Employment Opportunity Commission, Amicus brief in Mobley v. Workday, Inc., April 2024.
- U.S. District Court for the Northern District of California, Order Granting Preliminary Collective Certification, Mobley v. Workday, Inc., May 16, 2025.
- U.S. District Court for the Northern District of California, Order Granting in Part and Denying in Part Motion to Dismiss, Mobley v. Workday, Inc., ECF No. 360, June 22, 2026.