How to use LLMs in recruitment: a practical guide

Large Language Models (LLMs) are powerful models that soon no recruiter can live without, here's how to start using them.

How to use LLMs in recruitment: a practical guide

Disclosure: some links in this article are affiliate links. If you sign up through one, HeroHunt may earn a commission at no extra cost to you.

To effectively incorporate Large Language Models (LLMs) like ChatGPT, Claude or Gemini into the recruitment process, it’s essential to understand both the potential and the limitations of these technologies. 

This guide aims to provide a practical framework for leveraging LLMs in various stages of recruitment, from sourcing candidates to final interviews, while also emphasizing the importance of human oversight and ethical considerations. The short version: LLMs are excellent at producing language and unreliable at producing judgement. Almost every good use in recruiting sits on the first side of that line, and almost every expensive mistake comes from assuming it sits on the second.

An introduction to LLMs in recruitment

1. Understanding the capabilities of LLMs in recruitment

Before diving into the practical applications of LLMs in recruitment, it's crucial to grasp what these models can do. 

LLMs can process and generate natural language text, which allows them to perform tasks like resume screening, job matching, and even preliminary interviews. When LLMs are used in the context of actionable AI, AI that can execute on actions, we're talking about AI Agents.

LLMs are excellent at handling large volumes of data, identifying patterns, and automating repetitive tasks.

Key capabilities include:

  • Resume and cover letter analysis: LLMs can quickly scan through thousands of resumes and cover letters to identify candidates who meet specific job criteria.
  • Job description generation: By inputting a few key skills and qualifications, LLMs can help craft detailed and attractive job descriptions.
  • Candidate sourcing: LLMs can assist in finding candidates on professional networks by analyzing profiles and matching them to job requirements.
  • Automated initial screening: Through chatbots or automated emails, LLMs can conduct initial screenings to assess basic qualifications and interest levels.

That list is the promise, and it is worth stating the counterweight in the same breath, because the four items are not equally safe. Generating a job description is a task where the model produces a draft and a human immediately judges it, so an error costs you thirty seconds. Scoring a resume is a task where the model produces a number, the number looks authoritative, and nobody checks it. The first is a genuine productivity gain. The second is where the legal and quality risk in this entire field is concentrated, and the rest of this guide keeps coming back to that distinction.

2. Integrating LLMs into the recruitment workflow

The integration of LLMs into the recruitment process should be strategic and focused on enhancing efficiency and effectiveness. 

Here's how to do it:

  • Identify areas of need: Start by pinpointing the stages in your recruitment process that could benefit from automation or enhanced analysis. Common areas include candidate sourcing, initial screening, and communication.
  • Choose the right LLM tools: Not all LLMs are created equal. Select tools that are specifically designed for recruitment purposes and offer robust privacy and data protection features.
  • Give the model context, do not fine-tune it: This advice has reversed since 2023. Fine-tuning a model on your historical hiring decisions teaches it to reproduce those decisions, including the biased ones, which is exactly how Amazon’s famous resume-screening experiment failed. Feed context at query time instead: the job description, the scorecard, the CV.
  • Integrate with existing systems: Ensure that the LLM tools you choose can seamlessly integrate with your Applicant Tracking System (ATS), such as an AI-powered platform like Manatal, and other recruitment software to streamline the process.

The last two points matter more than they look. A recruiter copy-pasting CVs into a consumer chat window is running an unlogged, unauditable, and probably non-compliant screening process on the side of the official one, and it leaves no record you could produce if a rejected candidate ever asks how the decision was made. The same model reached through your ATS, reading structured fields, writing its output back to the candidate record, is the same technology with a defensible paper trail around it. Where the LLM sits in your stack is not an IT detail. It is most of the risk.

Highlight

Manatal

That makes the system of record a buying decision, and it is cheaper than most recruiters expect. Manatal is an ATS with AI candidate scoring, matching against a job description and generated job descriptions on its entry tier at $15 per user per month billed annually ($19 month to month), with a 14-day trial and no card required: Manatal pricing. Two caveats before you click. That tier caps at 15 active jobs and 10,000 candidates, so an agency desk carrying more open roles than that is really pricing the $35 tier, and API access sits higher again at $55. And its built-in scoring is a black box you cannot audit yourself, which is exactly the question to put to any vendor if you hire in New York City.

Start free on Manatal

Where LLMs actually help, stage by stage

The honest ranking of LLM use cases in recruiting tracks one variable: how quickly a human notices when the output is wrong. Rank your own use cases that way and you will get the same order.

Job descriptions, adverts and outreach copy. This is the strongest use case and it is not close. You supply the requirements, the model supplies the draft, you edit. Errors are visible instantly and cost nothing. Most recruiters who report real time savings from LLMs are saving it here, on the writing, not on the deciding.

Boolean strings and search syntax. LLMs are unreasonably good at turning a rambling description of a role into a working boolean string with the right synonyms, adjacent titles and exclusions. Ask for the string, run it, and the search results tell you immediately whether it worked. Again, fast feedback, low risk.

Summarising and structuring. Turning a 6-page CV into a 5-line summary against a scorecard, or a 45-minute interview transcript into structured notes, is a large real saving. The caveat is that summarisation is where hallucination actually shows up in recruiting: models will smooth over a career gap, merge two employers, or confidently state a tenure the CV does not support. Summaries are drafts to be checked against the source, not replacements for it.

Screening and scoring. Genuinely useful for triage on high-volume roles, genuinely dangerous as a decision. The section on bias below explains why in numbers.

Interviews. LLMs write good structured interview guides and good follow-up probes. Letting a model conduct and evaluate the interview is a different product with a different risk profile, and in the EU it is squarely a high-risk system.

What LLMs are still bad at in 2026

Model quality has improved enormously since this guide was first written, and three limitations have not moved at all, because they are structural rather than a matter of scale.

The first is that there is no ground truth for “good hire”. An LLM can tell you a CV matches a job description, which is a language-similarity question it is very good at. It cannot tell you the person will succeed in the role, because nothing in its training data connects a CV to an outcome two years later. Any tool claiming to predict performance from a resume is claiming something the underlying technology cannot do. Match quality and hire quality are different things, and vendors blur them constantly.

The second is inconsistency. Ask a model to score the same CV twice and you can get two different numbers. That is fine for brainstorming and disqualifying for a decision you may have to defend. If you use scoring at all, fix the temperature, fix the prompt, log the output, and treat the score as one input into a human decision rather than a gate.

The third is fluency masquerading as accuracy. The output is always well written, whether or not it is right. Recruiters are the target market for this failure mode because the artefact you receive, a confident paragraph about a candidate, looks exactly like a good one whether it is true or invented.

Prompts that actually work in recruitment

Most disappointing LLM output in recruiting is a prompt problem, and the fix is usually the same: stop asking for a verdict, start asking for evidence. The pattern below is worth more than any prompt library.

A bad screening prompt is “Rate this candidate out of 10 for this role.” It invites the model to invent a judgement it has no basis for, and you learn nothing about how it got there. A good one is “For each requirement in this job description, quote the exact line from the CV that provides evidence for it, or write NOT FOUND. Do not infer, do not summarise, quote.” The output is now checkable in seconds against the source document, the model cannot smuggle in a hunch, and you have a record of the reasoning rather than a number.

The same inversion works everywhere. Instead of “write outreach to this candidate”, try “list the three things in this profile that would make this specific role a step up for this person, then write a 90-word message that references exactly one of them.” Instead of “is this candidate a fit”, try “what are the three strongest reasons to reject this candidate, and what evidence would change your mind.” You are using the model for what it is good at, reading a document and producing language about it, and keeping the judgement where it belongs.

One more practical habit: give the model the scorecard, not the job advert. Job adverts are marketing documents full of unmeasurable adjectives. A scorecard with five concrete requirements produces dramatically better and more auditable output, and writing one is the sort of discipline that improves your hiring whether or not an LLM is involved.

What you can and cannot paste into an LLM

A CV is personal data, and under the GDPR the tool you paste it into is your processor. That single sentence rules out most of what recruiters actually do with LLMs today.

On consumer tiers of the major assistants, your conversations may be used to improve the models by default, with an opt-out buried in settings. Business tiers reverse that default: ChatGPT Business is $25 per user per month month to month, or $20 billed annually with a two-seat minimum, and workspace data is not used to train OpenAI’s models. That price difference between a free account and $20 a seat is the entire distance between an unlawful processing arrangement and a defensible one, and it is the cheapest compliance you will ever buy.

Beyond the training question, three rules keep recruiters out of trouble. Have a data processing agreement with whichever provider you use, which you cannot get on a personal account. Redact what the model does not need: it can assess a CV perfectly well without a name, an address, a photograph or a date of birth, and stripping those is also a cheap bias control. And tell candidates. Under the GDPR they have a right to know, and in New York City telling them is not optional.

Bias, and the law as it stands in 2026

This is the part of the original guide that said “ethical considerations” and left it there. It deserves specifics, because the evidence is now unambiguous and the law has moved.

University of Washington researchers ran more than 550 real CVs through three production LLMs, varying nothing but the name at the top. The models preferred white-associated names 85% of the time and male-associated names 52% of the time, and across roughly three million comparisons they almost never preferred a Black male name over a white male name - University of Washington. Nothing about the candidates changed. Only the name did.

The follow-up study is the one that should worry anyone relying on a human-in-the-loop as their safeguard. When 528 people screened candidates alongside a simulated biased AI, they largely reproduced the model’s bias in their own picks, and when it was neutral they were close to neutral - University of Washington. Human oversight is not a magic disinfectant. A recruiter reviewing a ranked list mostly ratifies the ranking. If you want oversight to mean anything, it has to happen before the model anchors the human, which in practice means blind-ish inputs and structured criteria rather than a reviewer nodding at a shortlist.

The law caught up in three places worth knowing. New York City’s Local Law 144 requires an annual independent bias audit of any automated employment decision tool used on NYC candidates, a published summary of the results, and at least 10 business days’ notice to the candidate, with penalties of $500 to $1,500 per day, per violation. Enforcement was long considered toothless, and that changed: the State Comptroller reviewed the same 32 employer disclosures the city had cleared and found at least 17 potential compliance failures against the city’s one, calling the enforcement ineffective - Office of the New York State Comptroller. Proactive investigation followed.

In Europe, recruitment and candidate selection are named as high-risk under Annex III of the AI Act. The obligations were due to bite on 2 August 2026, and the Digital Omnibus deferred them to 2 December 2027, adopted by the Council on 29 June 2026 - Council of the EU. Read that as breathing room, not a reprieve. The transparency and AI-literacy duties are already live, the GDPR never went anywhere, and systems bought in 2026 will still be running in 2028.

In the US, the vendor is no longer a shield. In Mobley v. Workday, a California federal court let discrimination claims proceed against the software provider itself on an agent theory, and the age-discrimination claim is running as a nationwide collective - Seyfarth Shaw. “The tool did it” is not a defence, and the tool may be liable alongside you.

A realistic first 30 days

The teams that get value from this move in a deliberately boring order, and the order is the whole trick: start where the model writes, end where the model reads, and never start where it decides.

In week one, get a business-tier account so the data question is closed, and use it only for language: adverts, outreach drafts, boolean strings, interview guides. You will find the savings immediately and you cannot do much harm. In week two, add summarisation against your own scorecards, using the quote-the-evidence prompt above, and spot-check every output against the source until you have a feel for how often it drifts. By week three you know where the model helps in your specific process, which is the point at which it is worth moving the work out of a chat window and into the system of record, so that outputs land on candidate records and someone can reconstruct a decision later.

Only then, if at all, consider automated scoring, and treat that as a procurement decision rather than a prompt: ask the vendor for the bias audit, ask what the score is computed from, and ask what happens when a candidate contests it. If you hire in NYC you need that audit published anyway. If the vendor cannot answer, you have learned something useful for free.

Two structural choices are worth making early. Keep a human decision at every reject, not just at every hire, since rejection is where the legal exposure lives and where automation is most tempting. And keep the log: prompt, model, output, and who decided what. Most teams do this through their ATS, some through a dedicated AI recruiter that sources and reaches out on autopilot while writing everything back to the record. Either is fine. A folder of chat transcripts is not.

Highlight

HeroHunt.ai

If the sourcing and outreach half of that log is where your week actually goes, the place to move the work into is not necessarily an ATS. HeroHunt.ai is an AI recruiter rather than a chat window: it searches across a billion public profiles, screens with language models against criteria you set, and runs personalized outreach, with the search, the message and the reply all landing on a candidate record you can reconstruct a decision from later. The honest limits are the ones this guide has argued throughout. It is a sourcing and outreach engine, not a pipeline manager, so most teams keep an ATS alongside it. And no model here or anywhere predicts who will succeed in the role, so keep a human decision on every reject.

Try HeroHunt.ai free

The bottom line

Use LLMs everywhere the output is language and a human checks it in seconds, which is most of a recruiter’s week: the adverts, the outreach, the search strings, the interview guides, the notes. That is a real and immediate saving, it carries almost no risk, and it is available to you this afternoon for the price of a business seat.

Be slow and deliberate everywhere the output is a judgement about a person. The models are demonstrably biased on nothing more than a name, a human reviewer tends to mirror rather than correct that bias, New York City is fining for it, the EU has recruitment on its high-risk list from December 2027, and a US court has already held that the vendor can be liable alongside the employer. None of that means do not use the technology. It means the question to ask about any LLM in your process is not “how much time does this save?” but “when this is wrong, who notices, and how fast?” Every recommendation in this guide falls out of that one question.

Written by Yuma Heymans (@yumahey), who built HeroHunt.ai and its AI Recruiter, and has spent the last few years finding out in production which parts of recruiting an LLM genuinely does better than a person, and which parts it only appears to.

Pricing, model behaviour and AI regulation in hiring are all moving quickly. Figures here were verified in July 2026 against primary sources, but check current terms before you buy or rely on them.