ChatGPT Candidate Screening: 2025 AI Recruitment Guide

AI-driven candidate screening now accounting for over 75% of initial applicant evaluations, this is how to adopt it.

ChatGPT Candidate Screening: 2025 AI Recruitment Guide

Disclosure: some links in this article are affiliate links. If you sign up through one, HeroHunt may earn a commission at no extra cost to you.

Most recruiters who say they "use AI for screening" are not using a screening product. They are pasting a CV into ChatGPT and asking it who to interview. That distinction is worth being blunt about, because the two things carry very different risks.

The adoption is real. In SHRM's 2025 Talent Trends survey of 2,040 HR professionals, 51% of organizations said they use AI to support recruiting, and 44% said they use it to screen resumes. Recruiting is the HR function where AI has landed hardest, and 89% of the people using it report time savings.

What has not arrived is evidence that these models are good at the specific job most people are using them for: deciding who is worth talking to. The published research points the other way, consistently and uncomfortably.

So this guide is not a hype piece. It covers what controlled studies actually found when researchers asked GPT-4 to rank candidates, the screening jobs ChatGPT genuinely does well, the prompt patterns that hold up, how to wire it in without leaking candidate data, and where the law stood in 2026 in New York, Illinois, Colorado and the EU.

The short version, before you read another word: ChatGPT is an excellent reader and a poor judge. Use it to extract, structure, summarize and draft. Do not use it to score, rank or reject.

The Evidence: What Happens When You Ask ChatGPT to Rank People

Two University of Washington studies are the ones to know. Neither is a vendor whitepaper, and both used the boring, rigorous method: change one thing about a CV, hold everything else identical, and see what the model does.

It penalizes disability signals, and explains itself with ableism

Kate Glazko and colleagues took a real 10-page CV and made copies that were identical except for four added disability-related credentials: a scholarship, an award, a DEI panel seat and a student organization membership. Six disabilities were implied (deafness, blindness, cerebral palsy, autism, depression and "disability" generally). They then asked GPT-4 to rank the enhanced CV against the control for a real student researcher opening.

Across 60 trials, GPT-4 ranked the enhanced CV first only a quarter of the time. Note what that means: the enhanced CV was the same CV plus extra honors. Adding achievements made the candidate rank worse.

Asked to justify itself, the model produced explicit ableism. Of a candidate implied to have depression, GPT-4 said the CV had "additional focus on DEI and personal challenges," which "detract from the core technical and research-oriented aspects of the role."

The researchers then built a custom GPT with written instructions on disability justice and DEI principles. It helped: the enhanced CVs then ranked first 37 times out of 60. But look at the residue. The autism CV still ranked first only 3 times out of 10, and the depression CV only twice, essentially unchanged from stock GPT-4. Prompting reduced the bias. It did not remove it.

Anonymizing the CV does not save you

Kyra Wilson and Aylin Caliskan ran the name experiment at scale: 120 first names associated with white and Black men and women, varied across resumes ranked against over 500 real job listings in nine occupations, for more than three million resume-to-job comparisons.

  • White-associated names were favored 85% of the time, against 9% for Black-associated names.
  • Male-associated names were preferred 52% of the time, against 11% for female-associated names.
  • The systems never preferred a Black male name over a white male name.

One honest caveat on that study, because it is often misreported: it tested three production language model retrieval systems (from Mistral AI, Salesforce and Contextual AI), not ChatGPT. It is evidence about the class of technology and the training data underneath it, not a measurement of OpenAI specifically. The full paper is worth reading if you are the person signing off on a tool.

The practical lesson is the one people skip. Stripping names does not fix this. The model infers identity from schools, cities, organizations and phrasing. Blind screening built on a language model is blind in name only.

It does not agree with itself

There is a structural problem underneath the bias problem. A chat model samples its output. Run the same CV through the same prompt on Tuesday and Thursday and you can get a different ranking, with different reasoning, for reasons that have nothing to do with the candidate.

For a marketing email, that is charming variety. For a hiring decision, it is fatal, because you cannot defend a decision you cannot reproduce. When a rejected candidate's lawyer asks why applicant A was advanced and applicant B was not, "the model said so, though it says something different now" is not an answer.

What falls out of this

One rule, and everything else in this guide follows from it: never let the model produce the shortlist. Let it produce evidence, and let a human produce the shortlist.

Where ChatGPT Genuinely Earns Its Place

None of the above makes ChatGPT useless in screening. It makes it useful for a narrower and less glamorous set of jobs than the vendors imply. Every item below shares one property: it produces something a human reads and can check against the source document, and none of them ends in a decision.

  • Extraction, not evaluation: turn 200 unstructured CVs into a structured table (years in the relevant role, technologies actually shipped, certification held yes or no, notice period). This is the highest-value use by a distance, and it is the one nobody talks about, because it is plumbing.
  • Rubric drafting: give it the job description and ask it to propose the five things that actually predict success in the role, then argue with it. The rubric is the deliverable, and a human owns it.
  • Screening question design: it is good at generating role-specific questions that are hard to bluff, and at proposing follow-ups based on a candidate's stated experience.
  • Interview note summarization: compress a 45-minute transcript into evidence against your rubric. Verifiable, because the transcript is right there.
  • Job description rewriting: tightening bloated requirement lists and cutting the jargon that suppresses application rates.
  • Candidate communication drafts: outreach, scheduling, rejection notes that read like a person wrote them. A human still sends them.

Notice what is absent: culture fit, personality inference, potential, "soft skills scoring." A CV contains no reliable signal about any of those. When you ask for them anyway, the model does not find signal, it pattern-matches on class and background markers and hands you a confident paragraph. That is precisely the mechanism the disability study caught in the act.

Prompt Patterns That Hold Up

If you are going to do this, do it in a way that survives scrutiny.

  • Rubric first, CV second. Define and freeze the criteria before the model sees a single candidate. Otherwise you are letting it invent the standard and apply it in the same breath.
  • Demand quotes. "For each requirement, quote the exact line from the CV that supports it, or write NOT FOUND." This single instruction kills most hallucination and turns verification into a ten-second job.
  • Ask for extraction, not scores. A 0-100 fit score is false precision that launders a guess into a number. Ask what is in the document.
  • One candidate per conversation. Ranking a batch in one context introduces order effects and lets earlier candidates anchor later ones.
  • Run it twice. If the two answers disagree, that disagreement is your reliability estimate, and you have just learned something important for free.
  • Never ask "should we interview this person". Ask "which of these five requirements does this CV provide evidence for, and what is the evidence."

A prompt that follows all of this looks less like magic and more like a form. That is the point.

The Data Problem: Which ChatGPT You Use Matters

This is the part that gets skipped, and it is the part that creates real liability.

Consumer ChatGPT and business ChatGPT have different default data policies. On the consumer tiers, conversations are used to improve the models unless you turn training off in settings. On ChatGPT Team, Enterprise and Edu, and on the API, business data is not used for training by default. OpenAI documents this on its enterprise privacy page.

So pasting a candidate's CV into your personal ChatGPT account is a materially different act from doing it in your company's Team workspace. In the first case you have taken someone else's personal data, sent it to a third party you have no contract with, and defaulted it into a training pipeline.

Under GDPR, a CV is personal data you are processing on a lawful basis, your privacy notice has to cover it, OpenAI becomes a processor needing a data processing agreement, and Article 22 is the provision to read carefully before you let anything automate a rejection. In practice:

  • Never paste candidate data into a personal account. Ever.
  • Use Team, Enterprise or the API, so the default is not training on your candidates.
  • If you cannot, turn off model training first, and understand that this fixes the data question and not the bias question.
  • Set a retention policy for anything you send, and know that "we deleted it from our side" is not the same as deletion at the processor.
  • Tell candidates. In several jurisdictions this stopped being a courtesy and became a requirement.

Or Skip the Plumbing Entirely

There is an unglamorous truth here: almost everything in the prompt section above is work that a modern applicant tracking system already does, inside your system of record, with an audit trail and a data processing agreement you already signed.

If you would rather buy than build, Manatal is the cheapest credible option with AI scoring built in, and its documentation is unusually specific about the mechanics: it extracts up to 10 criteria from a job description (150 words or longer), lets you weight required versus preferred criteria, needs at least one work or education entry on a profile before it will score at all, and gives a requirement-by-requirement justification rather than a bare number. Its own support docs still mark the feature as Beta with fair usage limits, which is more honesty than most of the category offers.

Highlight

Manatal

Since this is the buy-instead-of-build decision, here is the money side of it. Manatal publishes $15 per user per month billed annually ($19 monthly), and that entry tier stops at 15 active jobs and 10,000 candidates, so anyone running more than 15 open reqs is really pricing the $35 tier ($39 monthly). The API you would need to pipe data into your own OpenAI prompts appears only on the $55 tier ($59 monthly). Two caveats matter more here than the price. Scoring inside an ATS gives you a versioned, requirement-by-requirement record you can produce in a dispute, which a chat log cannot, but it does not make the scoring unbiased, and it does not move your Local Law 144 audit obligation onto the vendor. Mobley v. Workday is testing whether a vendor shares that liability at all. The employer's exposure was never in question. Treat its output the way this guide treats ChatGPT's: a sort order for a human, not a reject button. The 14-day trial takes no credit card.

Start free on Manatal

Bigger platforms have shipped the same thing. Just note who is on the other side of the AI hiring litigation below before you assume that buying instead of building transfers the risk. It does not.

If you took one thing from this article, take this: "the AI did it" is not a defense in any jurisdiction that matters.

New York City: Local Law 144

In force since January 2023 and enforced by the Department of Consumer and Worker Protection since July 2023. It covers any Automated Employment Decision Tool that substantially assists or replaces discretionary decision-making for candidates in NYC. If you are covered, you owe an independent annual bias audit, a public summary of its results on your careers page, and 10 business days' notice to candidates before you use it.

Penalties run $500 for a first violation and $500 to $1,500 for each subsequent one, and each day of continued non-compliant use can count separately.

Here is the trap. The law is about function, not about product category. A ChatGPT workflow that ranks applicants can be an AEDT. Almost nobody has commissioned an independent bias audit of their own prompt, and a prompt that produces a different order every run is not obviously auditable at all.

Illinois: HB 3773

Effective January 1, 2026. It amends the Illinois Human Rights Act to prohibit AI that has the effect of discriminating on protected grounds in recruitment and hiring, explicitly including the use of zip code as a proxy, and it requires notice to applicants and employees when AI is used in those decisions.

Colorado: quieter than expected

The Colorado AI Act (SB 24-205) was going to be the broad US framework. It did not survive contact with the legislature. Governor Polis signed SB 189 on May 14, 2026, pushing the effective date to January 1, 2027 and substantially scaling the law back: the duty of care against algorithmic discrimination, the deployer risk-management programs and the impact assessments came out, leaving a narrower disclosure and transparency regime.

EU AI Act: employment is high-risk, but the clock moved

AI used in recruitment and selection sits in Annex III, which makes it high-risk and pulls in the heavy obligations. The date everyone had circled was 2 August 2026. The Digital Omnibus changed that: after the European Parliament's endorsement on 16 June 2026 and the Council's final approval on 29 June 2026, the Annex III high-risk obligations shift to 2 December 2027.

Read that as more runway, not as a reprieve. GDPR did not move, national employment law did not move, and Article 22 has applied the whole time.

The United States, federally

Two cases tell you where this goes. The EEOC settled its first AI hiring discrimination suit against iTutorGroup for $365,000, after the company's software was programmed to auto-reject female applicants aged 55 and over and male applicants aged 60 and over. Crude, deliberate, easy to prove.

Mobley v. Workday is the harder and more important one. In May 2025 the Northern District of California granted preliminary certification of a nationwide ADEA collective, and roughly 14,000 people opted in by the March 2026 deadline. The theory being tested is that a screening vendor can be liable as an agent of the employers using it. Whichever way it lands, the employer's own exposure was never in question.

An Adoption Checklist That Will Not Get You Sued

  • Write down the rubric before you write the prompt, and have a human own it.
  • Use ChatGPT for extraction and summarization. Keep every advance and reject decision with a named person.
  • Never ask for a fit score, culture fit or personality inference.
  • Require quoted evidence from the source document for every claim the model makes.
  • Run on Team, Enterprise or the API, never a personal account, with a DPA in place.
  • Tell candidates you use AI, and what for. NYC, Illinois and the EU are converging on this.
  • Keep a log: prompt version, model version, date, output, human decision. If you cannot reconstruct a decision, you cannot defend it.
  • Test your own pipeline the way the UW researchers did. Take one real CV, vary one attribute, run it 20 times. It costs an afternoon and it tells you more than any vendor's fairness whitepaper.
  • Assume any workflow that ranks or filters is an AEDT until your counsel tells you otherwise.

What To Expect Next

The honest forecast is unexciting. Models will keep getting better at reading documents, which makes the extraction use case stronger every year. There is no sign they are getting better at judging people, because the bias is in the training data, which is the internet, which is us.

The regulatory direction is one-way: notice requirements, audit requirements and documentation requirements are spreading, and the Digital Omnibus delay is a scheduling change, not a change of direction. The teams that will be fine are the ones that can produce a rubric, a log and a named human decision-maker for any hire, which is the same answer that was correct before any of this existed.

Used narrowly, ChatGPT gives recruiters back the hours they spend reformatting other people's documents. That is genuinely worth having. Used to decide who is worth talking to, it reproduces the exact biases you hired it to escape, and it does so at a scale and speed no human panel could manage.

If you want scoring with a requirement-level audit trail instead of a chat log, this is the cheapest credible starting point. Check the 15-job cap on the entry tier against your req load first.

Start free on Manatal

Read the two UW studies. Write the rubric. Keep a human on the decision. Everything else is implementation detail.