Data scientist recruitment: a brief guide

Data scientists are some of the smartest people alive and hard to find. This guide helps you to find, assess and hire them.

Data scientist recruitment: a brief guide

Disclosure: some links in this article are affiliate links. If you sign up through one, HeroHunt may earn a commission at no extra cost to you.

Data scientists: who you’re talking to

Data scientists are increasingly sought after by companies. Especially tech companies employ data scientists to get the right data and make sense out of data. Some companies employ data scientists to improve the product they are offering and some companies need data scientists to fuel the organization’s data driven decision making process.

The role of data scientist is one where many different qualities come together. Every data scientist has to understand what data has to lead up to (the desired business outcome), use complex statistical concepts to create valid insights and use the right tools and code to extract, scrub and analyse data.

The core skills of a data scientist find their origins in domain knowledge, statistics and programming.

Domain knowledge

The ideal data scientist has a good understanding of the desired business outcomes. In every company the data they work with is different. A data scientist should understand the features of the data for the company in question in order to handle the data correctly and deploy the right algorithms.

Statistics

Data scientists need a good understanding of statistics. The models they deploy are based on math and specifically statistical concepts. Concepts and approaches like linear regression, the bell curve, central tendency, variability, variance and standard deviation should be a piece of cake for them.

Programming

Getting data, data scrubbing and data analysis requires the use of the right tools, Application Programming Interfaces (API’s) and in many cases custom code. In data science code is written in for example Python or R and has to integrate with back-end code that is written in other languages. Other tech skills include:

  • Programming languages like Scala, JavaScript, SQL, Spark, C, and C++
  • Libraries like scikit-learn, pandas, NumPy, Matplotlib
  • Data tools like Excel, Tableau, Hadoop, SAS

Different types of data scientists

Data science is a broad domain and there are many things to master within the field. Therefore in many organisations, especially larger organisations, there are usually specialized roles within a data science team that collectively work towards a shared goal.

These are the basic types of data scientists that companies need:

Data engineer

Data engineers are focussed on getting and preparing data. The goal of the data engineer is to prepare data so it can be used for further analysis and decision making. The most important activities to achieve this are data extraction, data consolidation and data cleansing (data scrubbing).

Data researcher

Data researchers are focussed on finding patterns in data, providing insights from the data to their team or customers and build analytics solutions so data rookies can use the insights derived from data.

Machine Learning expert

Machine Learning experts are specialised in learning models and algorithms. They research, build and test learning algorithms that are deployed in self learning products or for organisational purposes.

Next to the above mentioned roles there can be specializations like the data quality engineer, database administrator, data modeler, BI engineer or data architect.

Your sourcing mix to find the best data scientists

Data talent can be found across a variety of sources. Many recruiters would start their search on LinkedIn but there are niche platforms that match a lot better with the sought after data talent pool. Running that many channels by hand is the real cost of this approach, which is why some teams put an AI recruiter such as HeroHunt.ai across the search first and keep their own hours for the evidence review the rest of this guide describes.

Highlight

HeroHunt.ai

The whole argument of this guide is that no single platform holds the data science talent pool, so you end up repeating the same search on Kaggle, GitHub, Hugging Face, LinkedIn and Reddit. HeroHunt.ai is built for exactly that shape of problem: it searches across more than a billion profiles, screens them with language models against your brief instead of keyword matching, and writes the outreach. The honest caveat, and it matters on this particular role: it will not read a Kaggle leaderboard placement or a model card for you, and those artefacts are the signal this guide keeps telling you to trust. Use it to build and screen the list at scale, then judge the shortlist yourself on what they actually built.

Try HeroHunt.ai free

Kaggle

Kaggle is an online community for data scientists and machine learning experts and enthusiasts where Kagglers participate in data science challenges. On the platform users also share data sets, collaborate on code and solve data science challenges. Companies can post their challenges to Kaggle so users can choose to compete in the data challenges and have the opportunity to win prize money.

Kaggle has passed 18 million registered users according to its own milestone posts on the platform, and it has been owned by Google since 2017. Since 2023 it also hosts Models, so you can see who publishes and fine tunes models, not only who competes.

Between the competition leaderboards, the public notebooks and the models, Kaggle remains the single richest source of evidence about what a data scientist can actually do, rather than what they claim on a CV.

Kaggler profiles are very rich in relevant information about skills and activity with particular technologies, libraries and frameworks used. 

Here’s how to source Kaggle.

Stack Overflow

‍Stack Overflow is a question and answer website for engineers. Users can earn reputation points and "badges" by providing valuable answers. Next to the reputation of individual engineers you can also find a lot of information about the most recent technologies they have been working with.

Most of the information like top technology tags, reputation, badges and scores are based on actual activity rather than own input which makes the information very reliable from a sourcing perspective.

Read the dates before you read the reputation. Stack Overflow's question volume has collapsed since ChatGPT launched: from a peak of more than 200,000 questions a month in early 2014 to 3,862 questions in December 2025, a 78% drop year on year. The profiles and the reputation scores are still there and the skills history is still real, but for most engineers the activity signal is now years stale. Treat a Stack Overflow profile as a record of what someone worked with, not as evidence of what they are working on today.

Here’s how to source Stack Overflow.

GitHub

GitHub is a code repository and version control platform, fuelled by the functionality of Git, plus additional features. GitHub accounts are free and are frequently used to host open-source projects where engineers deposit their repositories. 

The benefit of sourcing on GitHub is that the information on talents is very up to date and relevant to their technical skills. If you are willing to take some time to research candidates on, you start to see which users are active and developing and sharing relevant code.

GitHub reported more than 180 million developers in its Octoverse 2025 report, with over 36 million joining in a single year, roughly one a second. It is the platform with the most active engineering users, and unlike Stack Overflow its activity signal is getting stronger rather than weaker, which makes it the first place to look for proof of current work.

Here’s how to source GitHub.

Hugging Face

Hugging Face did not exist as a sourcing channel when most recruiting playbooks were written, and it is now the obvious gap in them. If you are hiring machine learning engineers or applied scientists rather than analysts, this is where the work actually happens.

The platform hosts more than 2 million public models, over 500,000 public datasets and more than 1 million Spaces (small hosted apps), and it states that more than 50,000 organisations use it. Every model, dataset and Space has named contributors with a profile.

What makes it useful for sourcing is the same thing that makes GitHub useful: the evidence is the artefact. A model card tells you what someone trained, what it was trained on, how it was evaluated and how many people downloaded it. Download counts are public, so you can sort by impact rather than by self description. Someone whose fine tune has 500,000 downloads has demonstrated something no CV bullet can.

Practical approach: find a model or dataset close to your problem domain, open its contributors, and work outwards through the organisations they belong to. Profiles often link to GitHub, a personal site or Twitter/X, which is usually a faster route to contact details than the platform itself.

LinkedIn

LinkedIn is the most actively used professional platform in the world. Many recruiters rely on LinkedIn as their single source of candidates. Even though it has a lot of users, recruiters might not find their desired data talent here because there are not a lot of data science candidates that have complete and up to date profiles. In addition to that, competition is fierce on LinkedIn. That said, LinkedIn can still be a good source to include in your sourcing channels. 

If you don’t have a LinkedIn premium account or LinkedIn Recruiter seat you can learn here to source LinkedIn without premium features.

Talent networks

Talent networks are platforms with vetted talent that primarily provide candidates for contract positions. Many talent networks provide data scientists, AI experts and machine learning engineers. Talent networks typically charge a 15 - 25% surcharge for any contracted candidate.

Some examples of talent networks:

  • Toptal
  • Wellfound (this was AngelList Talent until November 2022, when it was spun out as an independent company and renamed)
  • Talent.io

The trade off is speed against margin. A network is the right call when you need a specialist for a defined project and the surcharge is cheaper than the time your own sourcing would take. It is the wrong call for a permanent hire you intend to build a team around.

Alternative platforms to source data talent

Medium

Medium is an online publishing platform for social journalism with bloggers sharing their views and knowledge. There is a great variety of topics covered by the bloggers on Medium and you can also find data talent that is sharing blogs about technical or more abstract topics. 

The information on Medium is rich because keywords can be found in the content of the articles that are written.

Here’s how to source Medium.

Reddit

Reddit is an overlooked platform for sourcing talent and it is one of the biggest online communities existing today. Be careful which number you believe, though. Reddit reported 126.8 million daily active uniques in Q1 2026, but only 52.0 million of those were logged in. The logged out half have no profile and no post history, so the population you can actually source is roughly the logged in one. Redditors engage in all kinds of comical discussions but data talent can also be found discussing data in subreddits like these:

  • r/datascience
  • r/dataisbeautiful
  • r/datasets
  • r/MachineLearning
  • r/learnmachinelearning
  • r/LanguageTechnology
  • r/deeplearning

Here’s how to source Reddit.

Getting their contact details

This is where most data science sourcing quietly dies, and it is the step the channel lists never cover. Kaggle, GitHub, Hugging Face, Medium and Reddit all hand you a username and no way to reach the human behind it. There is no InMail, there is often no email on the profile, and a cold DM on a community platform is the fastest way to get reported. Two approaches, in the order you should try them.

The free method: GitHub commit emails

Git records an author email in every commit, and on public repositories that metadata is public. Two ways to read it:

  • Add .patch to the end of any public commit URL. The From: header at the top contains the author's name and email.
  • Call api.github.com/repos/OWNER/REPO/commits and read commit.author.email on each entry. No authentication needed for public repos, and it returns a batch at a time.

The honest limitation: GitHub lets users hide their address, and when they do you get a placeholder like 49699333+username@users.noreply.github.com, which is undeliverable. We checked the most recent commits on scikit-learn, pandas and numpy while writing this: roughly half the author addresses were real and reachable (personal Gmail accounts and corporate addresses like jakevdp@google.com), and the rest were noreply placeholders. Half is a very good hit rate for a method that costs nothing, so try it before you spend credits.

Contact databases

For everyone else, and for the LinkedIn side of your sourcing mix, you are looking at a B2B contact database. These resolve a name plus a current employer into a work email and often a phone number. Apollo.io and Lusha are the two most commonly used by recruiting teams at this end of the market, and both sell credits rather than seats of unlimited lookups, so the question to ask is not what the tool costs but what a resolved email costs. Apollo.io is the one to start with, mostly because its free tier lets you measure your own hit rate on your own shortlist before you commit any money.

Highlight

Apollo.io

Every channel in this guide gives you a profile, not an inbox, and Apollo.io is the cheapest way to close that gap at volume. The free tier gives each seat 900 general credits per year (released monthly, so it is 75 a month, not 900), and Basic is $49 per seat per month billed annually, or $65 month to month, which grants 30,000 credits per seat per year up front. Read the credit types before you budget: email credits are metered separately from that general pool and are nominally unlimited on free, capped in practice at 10,000 a month from a verified corporate domain and only 100 a month from a personal one like Gmail. The honest caveat for this guide specifically: Apollo is keyed on company and job title, so it is strong on a data scientist whose current employer you can see on LinkedIn and weak on exactly the handle-only Kaggle, GitHub and Reddit profiles we just sent you to. Run the free commit trick above first, then spend credits only on the people it cannot resolve.

Try Apollo free

Putting the mix together

There is no single source for data talent, and the recruiters who do this well stop looking for one. The pattern that works:

  • Evidence first. Kaggle, GitHub and Hugging Face show you what someone built. That is worth more than any profile summary, and it gives you something specific to open a conversation with.
  • LinkedIn for coverage, not discovery. Use it to identify the employer and confirm the current role once another channel has surfaced the person.
  • Check the timestamp on every signal. The Stack Overflow collapse is the clearest lesson here: a reputation score that looked current in 2022 may be describing someone's 2016 self.
  • Solve contact before you scale sourcing. A list of 200 GitHub handles you cannot email is not a pipeline.

Once the mix is producing profiles, contact details become the bottleneck. Apollo's free tier gives each seat 900 general credits per year, with email credits metered separately, which is enough to test the workflow on a real shortlist before you pay for anything.

Try Apollo free