The buyer's guide to the three companies that sell human intelligence to AI labs: how each one prices, where each one wins, and what changed in 2026.
Mercor's annualized revenue doubled to $2 billion in four months, and the company is now negotiating a round at a $20 billion valuation - TechCrunch. Surge AI crossed $1.2 billion in revenue in 2024 without a dollar of venture capital, and Scale AI sold 49 percent of itself to Meta for $14.3 billion, then watched its biggest customers walk out the door. Three companies, three completely different business models, and together they collect most of what frontier labs and enterprises spend on human data.
But here is the problem: none of the three publishes a rate card, and the differences between them are not the differences their websites describe. Mercor is a labor marketplace that charges a margin on expert hours. Surge is a managed-quality shop that sells finished datasets to roughly a dozen labs. Scale is a platform company that now earns a guaranteed minimum from its largest shareholder while rebuilding its customer list around governments and enterprises. If you choose the wrong one, you will not find out from the sales deck. You will find out three months into a contract when the quality, the neutrality, or the security posture does not match the job.
This guide breaks down how each vendor makes money, the real price benchmarks buyers can use in negotiation, the security and legal record of each company after a year of breaches and lawsuits, and a decision framework for matching vendor to use case. It also covers the challengers eating at the edges (Handshake AI, micro1, Prolific, Invisible) and the increasingly common fourth option: recruiting your own expert bench with an AI recruiter such as HeroHunt.ai instead of renting one at a 30 to 40 percent margin. Everything here reflects late 2025 and 2026 information, because in this market a twelve-month-old fact is already wrong.
HeroHunt.ai
If what you actually need is a bench of credentialed experts you can re-use across projects, the cheapest route is often to recruit them yourself rather than pay a marketplace margin on every billed hour. HeroHunt.ai runs an AI Recruiter that searches over 1 billion profiles and handles the outreach on autopilot; the Starter plan is $149 per month for 3 open positions and Pro is $249 for 10, metered on positions rather than seats, with an 8-day free trial. The honest caveat: it finds and engages the people, it does not vet their work, run the task tooling, or carry the QA, contracts and payroll that Mercor, Surge and Scale bundle into their price. Budget for that layer yourself before you compare totals.
Contents
- The 2026 landscape in one page
- How each vendor actually makes money
- Mercor: the expert marketplace
- Surge AI: the quality house
- Scale AI: the platform after Meta
- Head-to-head on what buyers care about
- Security, legal and labor risk
- Pricing playbook: what to budget and how to negotiate
- Beyond the big three: challengers and the build-your-own option
- Decision framework: which vendor for which buyer
- Future outlook: agents, environments and the end of labeling
- Conclusion
1. The 2026 landscape in one page
The single most important thing to understand about this market is that the three names in the title are not three versions of the same business. They are three different answers to one question: how do you get expert human judgment into a model at scale? Mercor answers with a marketplace that recruits professionals, vets them with an AI interview, and bills their hours to labs at a margin. Surge AI answers with a small full-time team that designs tasks, runs a curated contractor pool, and delivers finished, quality-controlled data. Scale AI answers with a platform, a large operations organization, and now a corporate parent that owns nearly half of it. The right choice depends less on brand than on which of those three shapes fits the work you are buying.
The second thing to understand is that the market re-sorted itself in June 2025 and has not settled since. When Meta paid $14.3 billion for a 49 percent stake in Scale and hired its founder Alexandr Wang, Google, which had planned to spend about $200 million with Scale that year, moved to cut ties within days - CNBC. OpenAI and xAI followed. The human-feedback work those labs pulled did not disappear. It flowed to Surge, which was already the preferred vendor for Anthropic and Google's Gemini team, and to Mercor, whose expert-hour model was easier to spin up quickly. A year later that redistribution is visible in the revenue numbers, and it is the reason a buyer today faces a real three-way choice rather than a default.
The third thing to understand is scale, in the ordinary sense of the word. Leading labs now spend on the order of a billion dollars a year each on human-curated data, and the category is growing at more than fifty percent annually - Pebblous. That money is concentrated. Scale, Surge, Mercor and Handshake together capture over three-quarters of the industry's revenue - Dealroom. Concentration cuts both ways for a buyer: it means the big three have the operational depth to run a ten-thousand-task project on a two-week deadline, and it also means each of them depends heavily on a handful of frontier-lab accounts, which shapes how much attention a mid-size enterprise buyer will actually get.
The chart below puts the latest reported annualized revenue of the leading vendors side by side. Treat it as directional: Mercor and Handshake report gross figures that include what they pay out to contractors, Surge's figure is a late-2025 run-rate reported by third parties, and Scale's is the 2025 revenue it disclosed to Forbes. The mix of gross and net is itself a lesson about this market, which we return to in the pricing section.
Latest reported annualized revenue by vendor (USD millions)
What the chart says in practice is that momentum and size are now different rankings. Scale remains the best-known name and still runs the largest operations organization, but on revenue it has been passed by two companies that were a fraction of its size two years ago. Mercor's $2 billion run-rate is a gross number, so its actual net revenue is closer to a third of that after contractor payouts, but even on a net basis it has grown faster than any vendor in the category - Sacra. Surge's roughly $1.4 billion run-rate was reported by revenue trackers in late 2025 and the company itself does not publish financials - Latka. Scale's just under $1 billion for 2025, up from $870 million in 2024, came with a floor: Meta agreed to pay at least $450 million a year for five years - Forbes.
Why this matters for a buyer: the vendor with the most momentum is also the one that had the worst security year, the vendor with the best quality reputation is the smallest and most selective about who it works with, and the vendor with the deepest platform now has a shareholder that competes with most of its historical customers. Every strength on this list comes attached to a specific weakness, and the rest of this guide is about pricing those trade-offs correctly.
2. How each vendor actually makes money
Before comparing quality or speed, a buyer needs to understand the unit each vendor sells, because that unit determines everything downstream: how you will be quoted, where the margin hides, what you can negotiate, and what happens to your cost when your requirements change mid-project. None of the three publishes a price list, but each has a well-documented mechanism, and the mechanisms are different enough that a like-for-like quote is nearly impossible to obtain without doing some translation yourself.
Mercor sells expert hours at a margin. A lab or enterprise engages professionals through the platform, the professional logs hours, and Mercor bills the buyer a rate that includes its take. Buyer-side pricing is never published; every buyer surface on the site ends in a demo request, and the company confirms no figure - UsagePricing. Third-party analysts consistently estimate the take at roughly 30 to 35 percent, with contractors receiving 60 to 70 percent of gross billings - Sacra. A worked example that circulated in the Chinese tech press captures the mechanics: a lab paying about $150 per hour for a physician's time sees roughly $95 of that reach the physician - 36Kr. The practical consequence is that Mercor's cost to you is transparent in structure (hours times rate) even though the rate itself is negotiated, and that the way to lower it is to negotiate the margin, the expert tier, or the hours.
Surge AI sells finished data on a per-task or per-project basis. Its revenue comes from usage-based pricing per annotation task, plus project contracts for larger reinforcement-learning-from-human-feedback programs, with the price varying by task complexity and the domain expertise required - Sacra. You are not buying a person; you are buying a deliverable with a quality guarantee attached, and Surge's small full-time team designs the task, trains and monitors the contractors, and owns the output. Its contractor pay is correspondingly structured per unit of work rather than per hour, with workers earning in the range of 30 to 40 cents per working minute on many tasks, which is well above crowd rates and well below Mercor's expert rates. The per-task model means your cost scales with volume and difficulty, not with how long a particular person took, and that is the core reason labs with mature evaluation pipelines prefer it.
Scale AI sells platform access plus managed services, in two tiers. Enterprise contracts bundle the Scale Data Engine, the GenAI platform, dedicated operations support and service-level agreements, and require a sales conversation - Scale AI. A self-serve tier lets smaller teams pay as they go by credit card, with the first 1,000 labeling units and the first 10,000 curated images at no cost. Historically Scale's per-unit rates on simple tasks have been quoted in the low cents, around 2 cents per image and 6 cents per annotation, with multipliers for language, consensus and complexity - Sacra. For expert-level generative-AI work, Scale routes through its Outlier brand, where contributors are paid hourly, so an enterprise buying RLHF from Scale is really buying a blend of platform fee and hourly labor.
The table below translates those mechanisms into what a buyer will actually encounter.
| Vendor | Unit you buy | How quoted | Where the margin sits | Published price? |
|---|---|---|---|---|
| Mercor | Expert hours | Hourly rate per expert tier, plus platform fee | Take of ~30-35% on billed hours | No (advertised payouts only) |
| Surge AI | Completed tasks / datasets | Per task or per project, by complexity | Difference between per-task price and contractor pay | No |
| Scale AI | Platform + managed labeling | Enterprise contract, or self-serve pay-as-you-go | Platform fee plus labor markup | Self-serve free tier only |
The reason this matters is that the three models fail in different ways when your requirements move. If you discover mid-project that the task is harder than you scoped, Mercor's cost rises linearly with hours, Surge's per-task rate gets re-quoted, and Scale's enterprise contract may or may not absorb it depending on how the statement of work was written. If your volume drops, Mercor's cost drops with it, Surge may have minimums on a project engagement, and Scale's annual commitment does not care. And if you need to swap the expert profile, say from generalist annotators to board-certified physicians, only Mercor can do it inside the same commercial structure, because the expert profile is the product.
How to apply this: decide, before any vendor call, whether you are buying labor, deliverables, or infrastructure. Write that decision at the top of your RFP. Then ask each vendor to quote in your unit, not theirs. Mercor can quote per completed task if you insist; Surge can quote per hour on a pilot; Scale can quote per unit on self-serve. A vendor that refuses to translate is telling you something about how flexible the engagement will be.
3. Mercor: the expert marketplace
Mercor is the fastest-growing company in the category and, on the numbers, the most volatile. The San Francisco company crossed $1 billion in annualized revenue in February 2026 and $2 billion by June, up from roughly $500 million the previous September - Dealroom. It closed a $350 million Series C at a $10 billion valuation led by Felicis in October 2025, and by July 2026 was in talks to raise $500 million at $20 billion - Forbes. Between those two rounds it suffered the worst security incident the industry has seen, lost Meta as a customer for a period, was sued by five of its own contractors, and acquired a reinforcement-learning environments company. Understanding Mercor as a buyer means holding all of that at once.
The product is a marketplace with an AI front door. Professionals apply, complete a roughly twenty-minute AI-conducted video interview, and are scored and matched to projects; the company says its network now exceeds five million experts across law, medicine, finance, software and dozens of other domains - Mercor. The active pool is much smaller than the registered one. Around 30,000 contractors are weekly active, and payouts run above $2 million a day - Mercor. For a buyer, the relevant number is the second one, because it tells you how deep the bench really is when you ask for forty tax attorneys by Monday.
Mercor's growth playbook was a direct bet against Scale's. Where Scale built a crowd and then tried to move upmarket into expertise, Mercor started at the top of the labor market and recruited the kind of people labs could not otherwise reach: practicing physicians, investment-banking associates, litigators, PhDs. That positioning is why its founders were sued by Scale in September 2025, when Scale alleged that its former head of engagement management downloaded more than 100 customer strategy documents before joining Mercor as a general manager - TechCrunch. Mercor denied using any of the material, and its company profile at Sacra records the case as later voluntarily dismissed with prejudice. The episode matters less for its legal outcome than for what it reveals: by late 2025 Mercor was winning the same enterprise accounts Scale had built its business on.
How Mercor works for a buyer
The engagement model is closer to staffing than to data vending, and buyers who treat it as staffing get the best results. You define the expert profile, the task, the hours and the duration; Mercor sources and screens candidates, presents a slate, and handles contracts and weekly payment through its platform. The buyer typically supplies the tooling, the guidelines and the quality process, or uses Mercor's newer managed offerings for evaluation, agent training and off-the-shelf datasets - Mercor. The difference from Surge is stark: with Mercor you are usually running the project; with Surge the vendor is.
Pricing follows the labor. The advertised average contracted rate across Mercor's open roles was $109 per hour on 26 August 2026, down from $141 in early June, with individual roles spanning $60 to $250 per hour depending on domain and seniority - UsagePricing. Cybersecurity researchers and physicians sit at the top of that band, machine-learning engineers cover the whole range, and design, HR and accounting experts cluster near $70 to $80. Those are payouts to experts; what the buyer pays is that rate plus the platform's take, which is negotiated per account and never published. The trend in those advertised rates, which we chart in the pricing section, is one of the few live signals a buyer has about supply and demand in this market.
Where Mercor is strongest is speed to credentialed expertise and commercial flexibility. The company can stand up a domain-specific team in days, scale it up and down by the week, and let you change the expert profile mid-engagement without renegotiating the structure. It is also the vendor most willing to work with enterprises outside the frontier-lab club, and its Mercor Enterprise AI offering explicitly targets companies that want to capture their own workflows and train agents on them - Sacra. Where it is weakest is exactly the flip side: you carry more of the quality burden, the platform's own security controls were shown to be inadequate in March 2026, and a marketplace with five million registrants and thirty thousand weekly actives has, by construction, a long tail of people who passed the interview but will never see a task.
The APEX benchmarks and the pivot to agents
Mercor's most visible product outside its marketplace is the APEX family of benchmarks, and for a buyer they are a useful window into what the company's experts actually produce. APEX-Agents, launched in January 2026, measures whether frontier agents can complete long-horizon, cross-application tasks drawn from the daily work of investment-banking analysts, management consultants and corporate lawyers; the tasks were written by vice presidents, managing directors and managers with five to ten years at firms like Goldman Sachs, McKinsey and Cravath - Mercor. The headline finding was that frontier models completed fewer than 25 percent of tasks on a first attempt, and roughly 40 percent when given eight tries.
The chart below is Mercor's own leaderboard from the APEX-Agents launch. It shows single-attempt completion rates by model, with confidence intervals, and it is worth a close look because it demonstrates the quality of the expert-authored task set as much as the models.
APEX-Agents: first-attempt task completion by frontier model

Notice how tightly the top five cluster between 18 and 24 percent, and how wide the error bars are relative to the gaps between models. That is what a hard, expert-written benchmark looks like: it discriminates between capability tiers, not between marketing claims. Mercor has since shipped APEX-SWE for software engineering in March 2026 and, with the spend-management company Ramp, APEX-Accounting in July 2026, where the leading model scored 56.4 percent - Mercor. For a buyer evaluating Mercor's experts, these benchmarks are the portfolio.
The July 2026 acquisition of Deeptune shows where the company is heading. Deeptune builds reinforcement-learning environments, faithful recreations of enterprise applications such as spreadsheets and CRM systems, in which agents can practice tasks and be graded by verifiers; Andreessen Horowitz had led its $43 million Series A earlier in the year - Mercor. Combining an expert network that writes tasks and verifiers with software that hosts the environments is a direct move from selling labeled data to selling the training loop itself. Mercor followed it in September with a public guide to training a 397-billion-parameter knowledge-work agent with the open-source SkyRL framework - Mercor. Buyers who need agent training environments rather than static datasets should read that as Mercor declaring the category its own.
For a longer view of where the company is going, Brendan Foody's April 2026 talk at Stanford is the most complete public statement of the strategy. He spends the first third on why labs are shifting from labeled data to agentic data, and the second half on what that does to the labor market for experts.
Brendan Foody (Mercor): Agentic data and the future of AI
The talk is also useful for what it does not dwell on. There is very little about security, six weeks after the breach, and very little about the tail of registered experts who never get work. A buyer should bring both topics to the first call. If you are also weighing Mercor from the worker's side, or want the full mechanics of its interview and matching engine, HeroHunt has a dedicated deep dive in Mercor 2026: How It Works, Pay, Alternatives.
4. Surge AI: the quality house
Surge AI is the vendor the other two are measured against on quality, and the one a buyer is least likely to be able to hire on demand. Founded in 2020 by Edwin Chen, a former Google, Facebook and Twitter engineer, it reached $1.2 billion in revenue in 2024 with about 110 employees and no outside capital - Wikipedia. Chen owns roughly three-quarters of the company, which is why Forbes puts his net worth near $18 billion and lists him as the richest newcomer on its 2026 rankings - Forbes. Its customer list is short and heavy: Anthropic, OpenAI, Google, Microsoft, Meta and the U.S. Army, about a dozen frontier labs in total, and that concentration is both its moat and its constraint.
The company's operating philosophy is that quality is a product decision, not a workforce decision. Surge keeps its full-time headcount tiny, around 130 people by 2026, and puts them on task design, tooling and quality measurement rather than on account management - Sacra. Behind that team sits a contractor pool the company describes as elite, roughly 50,000 specialists it calls Surgers, who are recruited for writing ability, subject knowledge and judgment rather than for volume. Chen's own summary of the model, given on Lenny Rachitsky's podcast in December 2025, is that Surge reached a billion dollars with fewer than a hundred people by obsessing over quality and refusing to fundraise or do PR - Lenny's Newsletter. For a buyer that translates into a vendor that will say no to work it does not think it can do well, and that will not compete on price.
Surge's biggest single strategic gain came from someone else's decision. When Meta bought into Scale in June 2025, the labs that had relied on Scale for human feedback needed a neutral vendor with frontier-grade quality, and Surge was the obvious one. The company that had never raised money began exploring its first round within weeks, initially seeking up to $1 billion at a valuation above $15 billion - SiliconANGLE. By the end of July, Bloomberg reported talks at $25 billion - Bloomberg. As of this writing no closing has been confirmed publicly, which is entirely in character for a company that treats disclosure as optional.
What Surge sells and how it prices
A Surge engagement starts from the task, not the headcount. The company's standard products are a data-labeling platform with a Python SDK, RLHF tooling for live chat and transcript rating, red-teaming workflows, and quality-monitoring dashboards - Sacra. Buyers get a scoped deliverable: so many preference comparisons in this domain at this agreement threshold, so many adversarial prompts against this policy, so many expert-written rubrics. Pricing is usage-based per task, with project contracts for larger RLHF programs, and it moves with task complexity and required expertise. Contractor pay is structured per unit of work, in the neighborhood of 30 to 40 cents per working minute on many tasks, which works out to a level above crowd platforms but far below Mercor's professional rates.
That last comparison is the key to understanding when Surge is the right buy. If the task is a well-defined judgment that a strong generalist writer with domain familiarity can make in minutes, and you need hundreds of thousands of them at consistent quality, Surge's per-task economics beat both rivals. If the task requires a licensed specialist reasoning for an hour, you are paying Surge to find and manage that specialist, and Mercor's direct-hour model may be cheaper. Where Surge is rarely the right buy is at the bottom of the market. It does not compete for bulk image or LiDAR labeling, and it does not run a self-serve tier.
The company's public research output is the clearest evidence of its quality bar, and in 2026 it has become a marketing engine in its own right. In July, OpenAI cited Surge's GDP.pdf benchmark in the GPT-5.6 release - Surge AI. In August, Surge showed that a model post-trained on non-coding office work gained 5.8 points on SWE-Bench Pro, an argument that broad professional data transfers to narrow technical skill - Surge AI. And on 18 August it launched the Tuesday Work Index, a composite of eight of its benchmarks spanning chart reading, long-context instruction following, everyday judgment, writing, professional constraints, multimodal reasoning, agents in realistic companies and frontier mathematics - Surge AI.
The chart below is Surge's own Tuesday Work Index tracker from the end of August 2026. It plots every frontier model release of the year against the index, with the leading edge highlighted.
Surge AI Tuesday Work Index, January to August 2026

Look at the shape of the purple frontier line: a jump of roughly 14 points between February and March 2026, then a flattening near 69 by June. That plateau is what expert-built benchmarks are for. It tells labs, and by extension buyers, where models are actually stuck on professional work, and it is exactly the kind of signal that comes from a vendor whose contractors are strong enough to write tasks that models still fail. The leaders at launch were Claude Fable 5 at 66.8 and GPT 5.6 Sol at 66.7, with DeepSeek V4 Pro reaching 59.7 at a fraction of the run cost - Surge AI.
Surge's weaknesses are the mirror image of its strengths. It is selective about customers, and an enterprise outside the frontier-lab tier may struggle to get on its roadmap at all. Its transparency to buyers about labeler demographics is limited, which matters for safety and fairness evaluations that need reproducible cohorts - Prolific. Its revenue depends on about a dozen accounts, so a large customer departure would be felt immediately. And it is facing a California class action, filed in May 2025, alleging that annotators are misclassified as independent contractors and go unpaid for training time - Bloomberg Law. None of these is disqualifying, but a buyer should know that the vendor with the best quality story also has the least appetite for accommodation.
5. Scale AI: the platform after Meta
Scale AI entered 2025 as the category's default vendor and entered 2026 as a company rebuilding around a new owner, a new customer base and, since August, a new chief executive. The pivot point was 12 June 2025, when Meta agreed to pay $14.3 billion for 49 percent of the company, valuing it at about $29 billion, and hired founder Alexandr Wang to run its superintelligence effort - CNBC. Within days Google, Scale's largest customer, began moving its roughly $200 million of planned spend elsewhere, and OpenAI, Microsoft and xAI followed to varying degrees. The reason was not quality. It was that no frontier lab wanted its post-training data flowing through a company half-owned by a competitor.
The financial consequences arrived in stages. In July 2025, interim chief executive Jason Droege cut 200 full-time staff, about 14 percent of a 1,400-person workforce, and ended relationships with 500 contractors, writing in a memo that Scale had scaled its generative-AI capacity too quickly and built too many layers of bureaucracy - TechCrunch. For the full year, Scale told Forbes it booked just under $1 billion in revenue, up from $870 million in 2024, a respectable number that nonetheless represented a sharp deceleration from the trajectory the company had been on before the deal - Forbes. The same report disclosed the cushion: Meta committed to spend at least $450 million a year with Scale for five years, or more than half its annual AI data budget, whichever is less.
What Scale did with that cushion is the story a buyer needs to understand. Droege's own year-end account claims more than $1 billion in new business signed in 2025, nearly half in the fourth quarter, a data business that returned to profitability in the second half, and an applications business that more than doubled its revenue - Scale AI. The new customers are not labs. They are enterprises such as Mayo Clinic, BP and Allianz, U.S. defense contracts totaling around $200 million in the fourth quarter alone, and a rapidly growing international public-sector business across Qatar, the UAE, Saudi Arabia and the UK. Scale also reports delivering more than 150,000 hours of robotics and physical-AI data, a segment where neither Surge nor Mercor competes seriously. In short, Scale lost the frontier-lab RLHF market and replaced it with governments, regulated enterprises and physical AI.
Products, pricing and the new leadership
Scale's product line is broader than either rival's, which is its most durable advantage for buyers with heterogeneous needs. The Scale Data Engine covers collection, curation, annotation, RLHF and evaluation; the GenAI Platform covers fine-tuning, hosting and evaluating models on enterprise data; Donovan serves national-security customers; and the Outlier and Remotasks brands are the labor front ends for expert and crowd work respectively - Label Your Data. On top of that sits the research arm. In March 2026 Scale launched Scale Labs, extending its SEAL safety and evaluation lab into a broader research hub, with benchmarks including Humanity's Last Exam and SWE-Bench Pro that have been adopted in major model releases - Scale AI.
The launch image below is from that announcement. It is illustrative rather than a data chart, but it marks the moment Scale repositioned research as a front-line product rather than a side project.
Scale Labs launch, March 2026

The figure at the center of concentric rings is a fair visual metaphor for Scale's 2026 pitch: one platform, with evaluation at the core and enterprise, government and lab customers arranged around it. Whether that pitch lands depends on the person hired to deliver it. On 10 August 2026 Francis deSouza took over as chief executive, arriving from Google Cloud where he had been chief operating officer and president of security products, after eight years running the genomics company Illumina - Scale AI. That is an enterprise-sales and regulated-industry résumé, not a frontier-lab one, and it confirms where Scale intends to compete.
deSouza gave his first television interview as chief executive on 3 September 2026. It is short, and it is the clearest statement yet of how Scale describes itself to the customers it now wants.
Scale AI CEO Francis deSouza: We are headed for a multi-model world
Listen for the phrase "verify and trust" and the emphasis on a multi-model world. The pitch is that enterprises will run several models and need one vendor to evaluate, tune and govern them all, which is a platform argument rather than a data-labeling argument. Notice also what is absent: any claim to be the vendor of choice for frontier post-training, which was Scale's entire identity eighteen months earlier.
On pricing, Scale remains the only one of the three with any public entry point. The self-serve Data Engine gives the first 1,000 labeling units and the first 10,000 curated images free, then bills pay-as-you-go by card with no minimum, while enterprise contracts bundle quality guarantees, dedicated operations staff and access to the GenAI platform behind a demo request - Scale AI. Buyers using the self-serve tier will find pricing calculated per task with fixed and variable components and multipliers for project settings; buyers on enterprise contracts should expect an annual commitment and should negotiate the unit economics explicitly, because the bundle makes the labor markup hard to see.
Where Scale wins is breadth, physical-AI and multimodal capability, government-grade compliance and the ability to absorb a very large, very messy project. Where it loses is neutrality, which is now structural rather than reputational, and the fact that its most experienced generative-AI operations staff were the ones cut in July 2025. A frontier lab or a company that competes with Meta should treat the 49 percent stake as a hard constraint. A hospital system, an energy major or a defense ministry should treat it as irrelevant and evaluate Scale on the strength of its platform, which is considerable.
6. Head-to-head on what buyers care about
Comparisons of these three vendors usually collapse into "Mercor for experts, Surge for quality, Scale for scale," and that shorthand is roughly right but not useful for a purchase decision. A buyer needs to know how each vendor performs on the six dimensions that actually determine whether a project succeeds: quality of judgment, speed to staff, cost structure, neutrality, security and compliance, and breadth of modality. On some of those the three are close; on others the gap is decisive.
The most important dimension, quality, is also the hardest to compare fairly because each vendor defines it differently. Surge defines quality as agreement, calibration and writing standard on a well-designed task, and it controls the whole chain to achieve it. Mercor defines quality as credentials and domain depth, and it leaves much of the task-level control to the buyer. Scale defines quality as consistency at volume, with consensus mechanisms and multipliers baked into its platform. For frontier RLHF and preference data, Surge is the consensus leader among the labs that can afford it. For work that needs a licensed specialist's reasoning, Mercor's bench is deeper. For multimodal, robotics and defense data, Scale has capabilities the other two simply do not offer.
The table below scores each vendor on the six dimensions as they stand in September 2026, drawing on the evidence in the preceding sections.
| Dimension | Mercor | Surge AI | Scale AI |
|---|---|---|---|
| Expert depth | Strongest (licensed professionals, ~30k weekly active) | Strong generalists and specialists (~50k Surgers) | Broad via Outlier; thinner at the top |
| Task-level quality control | Buyer-led unless managed service | Vendor-led, benchmark-grade | Platform-led with consensus tooling |
| Speed to staff | Days | Weeks; selective on customers | Days on self-serve; weeks on enterprise |
| Cost structure | Hourly plus ~30-35% take | Per task / per project | Platform fee plus labor; annual commit |
| Neutrality | Independent; VC-backed | Independent; founder-controlled | Meta owns 49% |
| Security and legal record | Major breach Mar 2026; contractor suits | Misclassification class action 2025 | Outlier wage suits 2024-25; government-grade compliance |
Reading down the neutrality row explains most of the last year's customer movements. Surge is majority-owned by its founder, Mercor by a syndicate of venture investors with no model-building interests, and Scale by a company that trains its own frontier models. For a lab, that row alone decides the question. For an enterprise that does not compete with Meta, it is close to irrelevant, and the breadth row matters more.
Reading across the security row is where a buyer's legal team will spend its time. Mercor's March 2026 breach exposed data on tens of thousands of contractors along with source code and API keys, and Meta's response was to pause work indefinitely - The Next Web. Surge's exposure is a labor-classification case rather than a data incident. Scale's is a series of wage-and-conditions suits from Outlier contractors, offset by the fact that it holds the compliance certifications government buyers require. Each of those is a different kind of risk, and the next section treats them in the detail a procurement review needs.
Why this matters: the head-to-head shows there is no dominant vendor, only a dominant vendor per use case. How to apply it: score your own project on the same six rows before the vendor calls, weight the rows by what would actually kill the project, and let the weighted score, not the brand, pick the shortlist. Most buyers who do this honestly end up with two vendors, not one, and a plan to split work between them.
7. Security, legal and labor risk
The human-data industry had a bad year on trust, and a buyer signing a contract in late 2026 is signing it in that context. Three separate categories of risk have surfaced: supply-chain security, worker classification, and inter-vendor litigation. Each of the three vendors has been touched by at least one, and the way each responded tells you what to expect if something goes wrong on your project. This section is deliberately blunt because the sales conversations will not be.
The largest single event was the Mercor breach. On 27 March 2026 a threat group compromised the release pipeline of LiteLLM, a widely used open-source library for connecting applications to AI services, and published malicious versions that harvested credentials for roughly forty minutes before removal; Mercor confirmed on 31 March that it had been affected and attackers claimed about 4 terabytes of data including candidate profiles, personal information, employer data, source code and API keys - TechCrunch. Meta paused all work indefinitely, OpenAI opened an investigation, and five contractors filed suit. Mercor's June update says it brought in Google's Mandiant and Latacora, rotated credentials across every cloud, code and SaaS system, deployed tighter network controls and 24/7 managed detection, and that "all frontier labs have increased their work with us over the last few months" - Mercor. Both halves of that are true and both matter: the remediation is real, and the revenue recovered, but the incident showed that a vendor's own dependencies are your attack surface.
Worker classification is the second category, and it affects Surge and Scale more directly than Mercor. Surge was sued in San Francisco Superior Court in May 2025 by annotators alleging that they were misclassified as contractors and denied minimum wage, overtime and paid training, with the complaint citing millions in unpaid wages - Inc.. Scale faced two suits within a month at the turn of 2025: one Outlier contractor alleged he was promised $25 an hour and paid a fraction of it, and a second suit alleged that contractors were required to review disturbing content without adequate protection - TechCrunch. Mercor's higher-paid professionals are less exposed to wage claims, but its breach lawsuits are a reminder that a marketplace holds the personal data and video interviews of everyone who ever applied.
The third category is vendors suing each other, and it is worth understanding because it reveals how much customer knowledge moves between these firms. Scale's September 2025 suit against Mercor and its former engagement head alleged that more than a hundred customer strategy documents were copied to a personal drive before the employee joined Mercor; Mercor's co-founder Surya Midha replied that the company had no interest in Scale's secrets and had offered to have the files destroyed before the suit was filed - Axios. Whatever the merits, the lesson for a buyer is that your project details, pricing and roadmap are exactly the kind of material that walks between vendors with people. Confidentiality terms need to be written with that in mind.
A less-noticed thread in the Mercor incident is what it revealed about compliance paperwork. The compromised library had been certified by Delve, an AI compliance startup that was subsequently accused of faking security-certification data, after which Y Combinator severed ties with it and LiteLLM switched compliance partners - TechCrunch. For a buyer, the lesson is that a vendor's SOC 2 badge tells you about the vendor, not about the forty open-source packages the vendor's engineers installed last quarter. The questions that would have surfaced the risk are about dependency management and secrets rotation, not about certifications, and they belong in the security questionnaire alongside the usual attestations.
The commercial mechanism for pricing these risks is indemnity, and here the three vendors differ in what they will sign. A marketplace like Mercor is structurally reluctant to warrant the employment status of five million registrants, so buyers typically get a compliance representation plus a capped indemnity. Surge, with its smaller and more controlled pool, has more room to warrant but less commercial need to concede. Scale, selling to governments, already carries the contractual apparatus and will usually accept broader terms in exchange for the annual commitment. The negotiation is easier when you know which of those postures you are facing before the redline arrives.
Given all of that, a buyer's contract should carry a small number of non-negotiable provisions, which the three vendors will accept with varying degrees of enthusiasm.
- Dependency and supply-chain disclosure: a list of third-party libraries and services that touch your data, and notification within 24 hours of any compromise upstream
- Data segregation and deletion: your prompts, outputs and guidelines in isolated storage, with certified deletion at project end and no re-use for other customers
- Conflict-of-interest clause: named restrictions on staff who work your project also working for named competitors, and disclosure of any shareholder with a model-building business
Two further provisions round out the set and are worth writing in prose because vendors argue about the wording. A workforce status warranty has the vendor confirm that its contractor arrangements comply with applicable labor law and indemnify you against classification claims, which matters because the Surge and Outlier suits make it likely that customer names appear in discovery. Audit and incident rights give you the ability to review security attestations, penetration-test summaries and incident reports on request rather than waiting for a press story. Each of the five has a real-world trigger in the last eighteen months: the dependency clause exists because of LiteLLM, the segregation clause because Meta's stake made co-mingled data a competitive problem overnight, and the conflict clause because of the Scale-Mercor litigation. A vendor that pushes back hard on any of them is telling you which risk it is not prepared to own, and that is useful information before signature rather than after.
8. Pricing playbook: what to budget and how to negotiate
Because none of the three vendors publishes a rate card, the buyer's pricing task is to build a benchmark from the pieces that are public, then negotiate against it. The pieces are more available than they look. Mercor advertises what it pays experts. Scale publishes its self-serve free tier and third parties have documented its historical per-unit rates. Surge's contractor pay and pricing model are documented by analysts even though its buyer prices are not. Assembled correctly, these give a buyer a floor, a ceiling and a target for each vendor.
Start with labor cost, because it is the largest component for all three. On Mercor, the average advertised contracted rate across open roles moved from $141 per hour in early June 2026 to $109 by late August, a decline of more than twenty percent in eleven weeks - UsagePricing. Part of that is mix, as more mid-tier roles were posted, and part is a genuine softening as supply of vetted experts caught up with demand. Either way it is the most transparent live price signal in the category, and the chart below tracks it. Note that the tracker relabeled the metric at the end of June, which is why the series starts in early June and skips that week.
Mercor advertised average contracted expert rate, summer 2026
The practical reading of that line is that a buyer negotiating a Mercor engagement in September 2026 should not accept expert-rate assumptions built on spring 2026 numbers. If a proposal prices physicians at the top of the $110 to $250 band, ask for the distribution of rates actually paid on comparable projects in the last sixty days. A falling market is a negotiating asset only if you know it is falling. Add the platform's take on top: at a 30 to 35 percent margin, a $109 expert hour costs the buyer roughly $150 to $165, and a $200 specialist hour lands near $275 to $300.
Surge's arithmetic runs the other way, from the task up. If a task takes a Surger ten minutes and Surge pays 30 to 40 cents per working minute, labor cost per task sits around $3 to $4, and the buyer price will be a multiple of that reflecting task design, QA, tooling and margin. A buyer can back into Surge's effective hourly rate by dividing the quoted per-task price by the expected minutes per task, and should do so on the pilot, because that number, not the per-task price, is what to compare against Mercor. For frontier-lab RLHF programs, Surge's project contracts run into the tens of millions per year for its largest accounts, which is consistent with a billion-dollar business concentrated on about a dozen customers - Sacra.
Scale's arithmetic depends on tier. On self-serve, the free allocation of 1,000 labeling units lets a team benchmark quality before spending anything, after which per-unit pricing applies with multipliers for project settings - Scale AI. On enterprise, the useful public anchor is that Scale's largest single customer relationship, Meta, carries a floor of $450 million a year, and that Google's planned 2025 spend was about $200 million; a mid-size enterprise buying a bundled platform-plus-labeling contract should expect to be quoted in the high six to low seven figures annually with an annual commitment. The self-serve tier is the best free due-diligence tool in the category and there is no reason not to use it, even if you intend to buy enterprise.
A worked example makes the translation concrete. Suppose a buyer has $1 million for a six-month program of expert evaluation in a professional domain. On Mercor at an all-in rate near $160 per hour, that is roughly 6,250 expert hours, or about ten full-time specialists for the period, with the buyer supplying tooling and QA. On Surge, if the same judgments can be decomposed into fifteen-minute tasks priced in the low tens of dollars each, the budget buys on the order of 40,000 to 60,000 finished, quality-controlled tasks and no internal QA headcount. On Scale's enterprise tier, the same money buys platform access, a managed labeling program and evaluation tooling, but a meaningful share goes to the platform rather than to labor. None of those is obviously the better deal; which one is depends on whether the constraint is expertise, volume or infrastructure.
With those benchmarks in hand, three negotiating levers work across all three vendors.
- Quote in your unit: insist on a per-task or per-hour figure alongside the vendor's preferred unit, so the three proposals can be compared
- Separate labor from platform: ask each vendor to break out expert or contractor cost from fees, tooling and QA, which exposes the margin
- Pilot with a quality gate: fund a two-week pilot with an agreed acceptance rate, and tie the production price to hitting it
Two further levers are about timing rather than structure. Buy flexibility explicitly, meaning the right to change expert tier, volume and task type mid-contract at prices agreed in advance, because every vendor will price it before signature and almost none will grant it for free afterward. And use the self-serve tiers as leverage: run the same sample through Scale's free units and a small Mercor engagement to get real quality data before the enterprise conversation. The reason these levers work is that each vendor's weakness is the other vendors' strength. Mercor will separate labor from platform because that is how it is built; Surge will resist, so its willingness to do so on a pilot is itself a signal. Scale will happily run a self-serve pilot because it wants the enterprise deal, and the quality data from that pilot is exactly what to put in front of Surge. The buyer who arrives with a benchmark, a unit and a pilot plan gets a different conversation from the one who arrives asking what it costs.
9. Beyond the big three: challengers and the build-your-own option
The three vendors in the title are not the whole market, and in 2026 the alternatives matter more than they did a year ago. Two challengers have reached the scale where they belong on any serious shortlist, several specialists win specific use cases outright, and a fourth path, recruiting your own expert bench directly, has become practical enough that some labs and most enterprises should at least price it. This section covers all three, with the caveat that the challengers' numbers move fast and should be re-checked before purchase.
The largest challenger is Handshake AI, the human-data arm of the college-recruiting network. It reached nearly $1 billion in gross annualized AI-training revenue by April 2026, up from $5 to $10 million at launch fifteen months earlier, and nets roughly $300 million after paying contractors, who earn $100 to $125 an hour for specialized work in mathematics, physics and computer science - Sacra. Its distinctive asset is a pipeline of graduate students and early-career specialists that neither Mercor nor Surge can match on volume, and its economics look like Mercor's: gross billings, a take, and hourly experts. For buyers whose work needs deep STEM knowledge rather than professional licenses, Handshake competes directly with Mercor on price and increasingly on quality.
The second challenger is micro1, an AI-vetted talent platform that grew its gross annualized revenue to about $500 million in 2026 from $100 million earlier in the year - Dealroom. It is a direct Mercor analogue with a stronger emphasis on software engineers and a faster-moving, lower-priced bench. Beyond those two, Prolific wins where research-grade demographic control matters, with more than 300 filters and a verified participant pool across dozens of countries, which is exactly the capability it criticizes Mercor and Surge for lacking - Prolific. Invisible, at about $134 million of 2024 revenue, is the managed-operations option for enterprises that want a vendor to run the whole workflow. Turing and Toloka round out the list for engineering-heavy and multilingual work respectively. HeroHunt maintains a fuller ranking in Top 10 Human Data Providers in 2026.
The build-your-own option
The fourth path deserves more attention than it usually gets, because the marketplace economics make it compelling for a specific kind of buyer. Every dollar of expert time bought through Mercor, Handshake or micro1 carries a take of roughly a third; every task bought through Surge carries the cost of a management layer you may already have. If your organization needs the same fifty specialists for eighteen months, the arithmetic of renting them at a margin versus recruiting them once is not close. The traditional objection was that finding and engaging credentialed experts at scale was itself a full-time operation, and that objection has weakened as AI recruiters have matured.
This is where a tool like HeroHunt.ai fits into a human-data strategy. Its AI Recruiter searches more than 1 billion profiles, screens against a natural-language brief, and runs personalized outreach on autopilot; the plans are metered on open positions per month rather than seats, from $149 for three positions to $499 for twenty - HeroHunt.ai. For a data-operations lead who needs forty board-certified radiologists or two hundred senior Java engineers with a specific stack, that replaces the sourcing half of what a marketplace charges for. It does not replace the other half. Vetting work samples, running task tooling, handling contracts and payroll, and measuring quality remain your job, and that is precisely the layer Surge and Scale sell. HeroHunt's guides to assessing human data labelers and recruiting annotation talent for AI labs cover that layer in detail.
The buyers for whom build-your-own works best share three traits. They have durable, repeat demand for a definable expert profile, so the up-front recruiting cost amortizes. They already run or can cheaply run task tooling, which in 2026 usually means an open-source labeling stack plus a spreadsheet of rubrics. And they care about data confidentiality enough that keeping experts under direct contract, rather than on a shared marketplace, is worth the administrative load. Frontier labs meet all three, which is why most of them already run direct expert programs alongside their vendor spend. Regulated enterprises often meet the third and can be helped to meet the first two. Early-stage companies with spiky demand usually do not, and should rent.
The most common mistake in this decision is comparing a marketplace's all-in hourly rate against a recruiter's subscription and concluding that building is nearly free. It is not. A realistic build budget includes recruiting, a paid work-sample assessment for every candidate, contracting and payments infrastructure, a quality lead, and the tooling. For fifty experts, that is a real program with a real headcount. The right comparison is that program's annual cost against the marketplace margin on the same fifty experts' hours, and for durable demand the program usually wins by a wide margin. For everything else, the marketplaces exist for a reason. HeroHunt's labs guide to hiring AI-training talent walks through the build budget line by line.
10. Decision framework: which vendor for which buyer
Everything above reduces to a small number of questions, and the honest answer to each one points at a vendor. The framework below is built for a non-technical buyer who needs to defend the choice to a finance or legal reviewer, so it starts from what you are buying rather than from what the vendors sell. Answer the first question first; most of the decision is made there.
The first question is what the work actually is. If it is preference data, evaluation, red-teaming or rubric-writing at frontier quality, with a mature internal team to receive it, you are buying finished data and Surge is the default. If it is hours of a licensed or credentialed professional applied to tasks you will design and manage, you are buying labor and Mercor, Handshake or micro1 are the defaults, with Mercor strongest at the top of the market. If it is multimodal, robotics, defense or a large enterprise deployment that needs one platform for data, tuning and evaluation, you are buying infrastructure and Scale is the default. If it is a durable bench of the same experts for a year or more, you are buying a recruiting program and should price building it.
The second question is who you are, because it changes which risks bind. A frontier lab or any company competing with Meta should treat Scale's ownership as disqualifying for post-training work and choose between Surge and Mercor on quality versus flexibility. A regulated enterprise or public-sector buyer should weight compliance certifications and platform breadth, which favors Scale, and treat Mercor's breach as a due-diligence item rather than a veto. A research organization that needs reproducible demographics should look past all three to Prolific. The diagram below compresses those two questions into one path.
Follow the path that matches your work and then stop at the check node, because that is where the deal succeeds or fails. The Surge path fails if you are not a customer Surge wants, in which case no amount of budget helps and Mercor's managed services are the fallback. The Mercor path fails if your legal team cannot get comfortable with the vendor's security posture after March 2026, or if you do not actually have the internal capacity to run quality on expert output. The Scale path fails if your board, your customers or your own model strategy make Meta's stake a problem. The recruiting path fails if your demand is not durable enough to amortize the program.
Three worked scenarios show the framework in use. A Series B agent company building an evaluation suite for legal workflows has spiky demand, no internal QA team, and a confidentiality concern: Mercor's managed evaluation service for the first six months, moving to a direct bench of ten attorneys recruited with an AI recruiter once the rubrics stabilize. A hospital network fine-tuning a clinical documentation model has regulated data, a need for platform-level audit trails, and no competitive exposure to Meta: Scale's enterprise tier, with a self-serve pilot first to establish quality on de-identified samples. A frontier lab expanding post-training into professional domains has the internal team, the budget and the neutrality requirement: Surge for preference and evaluation data, Mercor for hourly specialists on the hardest tasks, and a direct program for the twenty or thirty experts it wants to keep.
Why this matters: each of those scenarios would have been mis-served by a single default vendor, and each would have been over-served by trying to use all three. How to apply it: run your project through the diagram, write down the check-node answer, and take that answer into the vendor call as your opening position rather than as a discovery.
11. Future outlook: agents, environments and the end of labeling
The category these three companies compete in is being redefined underneath them, and a buyer signing a multi-year contract should understand the direction. Three shifts are already visible in 2026: the unit of purchase is moving from labeled examples to agent training environments, AI is taking over the bottom of the labeling market, and the labs are pulling more expert work in-house. Each shift favors a different vendor, and the vendors' own 2026 moves show they know it.
The move to environments is the clearest. Mercor's acquisition of Deeptune was explicitly a bet that reinforcement-learning environments, simulated workplaces where an agent practices and is graded, would become "the defining bottleneck," and that the scarce input is experts who can write the tasks and verifiers - Mercor. Surge's Tuesday Work Index and its CoreCraft benchmark of agents operating inside realistic companies are the same idea from the evaluation side. Scale's research agenda under Scale Labs lists agentic and multimodal systems and long-horizon augmented workflows as its first priorities - Scale AI. When all three vendors reorganize around the same object, the buyer should assume that static datasets are becoming a commodity and that the premium will attach to environments, verifiers and expert-written rubrics.
The second shift is automation of the low end. Simple image, text and audio labeling is increasingly done by models with human review of the residue, which erodes the volume business Scale was built on and explains why its data business shrank before it returned to profit. That erosion does not reach the top of the market. The reason Mercor's APEX-Agents leaderboard tops out near 24 percent on first attempts, and Surge's index plateaus near 69, is that professional judgment tasks are exactly where models still fail, and every point of improvement requires more expert data, not less. The vendors with the deepest expert benches are structurally protected; the vendors with the largest crowds are not, which is the single best explanation for Scale's pivot to enterprise, government and robotics, where human data is still scarce and physical.
The third shift is disintermediation, and it is the one buyers control. Frontier labs now spend enough on human data, on the order of a billion dollars a year each, that recruiting their own expert programs makes obvious financial sense, and most of them have - Pebblous. What has changed in 2026 is that the tooling to do this is available to organizations far smaller than a lab. An AI recruiter can source a defined expert profile at volume, an open-source labeling stack can host the tasks, and the marketplaces themselves increasingly sell the managed layer separately. The likely equilibrium is not that Mercor, Surge and Scale disappear, but that they compete with their own customers' internal programs for the most durable work and keep the spiky, specialist and high-volume work that no internal program can justify.
For a buyer, the outlook translates into three contractual instincts. Prefer shorter terms and explicit flexibility, because the unit you are buying may change within the contract. Weight a vendor's environment and evaluation capabilities, not just its labeling throughput, because that is where the next two years of spend are going. And treat every vendor relationship as a bridge to a partially owned expert bench, because the economics of a 30 to 35 percent margin on durable demand will not survive contact with a competent recruiting program. Yuma Heymans (@yumahey), who built HeroHunt.ai and has spent years watching labs decide between renting expert workforces and recruiting their own, wrote this guide from the recruiting side of that line, and the view from there is that the vendors' most valuable customers are already becoming their competitors.
12. Conclusion
Mercor, Surge AI and Scale AI are not three brands of the same product, and the fastest way to waste a human-data budget in 2026 is to treat them as if they were. Mercor sells credentialed hours at a margin, grows faster than anyone, and had the industry's worst security year. Surge sells finished, benchmark-grade data to a dozen labs, prices per task, and does not need your business. Scale sells a platform with the broadest capabilities in the category, backed by a five-year revenue floor from a shareholder that most frontier labs consider a competitor. Each is the right answer to a specific question and the wrong answer to the other two.
The decision framework is short. Buy finished frontier-quality data from Surge if you can get on its roadmap and have the team to receive it. Buy expert hours from Mercor, or from Handshake and micro1 at the STEM and engineering end, when you will run the task and need the profile to flex. Buy platform and managed services from Scale when the work is multimodal, physical, regulated or governmental and Meta's stake does not bind. And when demand for a defined expert profile is durable, price a direct program with an AI recruiter, because the marketplace margin on a year of the same fifty experts pays for the program several times over.
Whichever path you take, three disciplines protect the budget. Quote in your unit, not the vendor's. Pilot with a quality gate before committing. And put the security, workforce and conflict clauses in the contract before signature, because every one of them has a real trigger in the last eighteen months. The vendors in this guide are all good at what they do. The buyer's job is to be equally clear about what that is.
Yuma Heymans (@yumahey) built HeroHunt.ai, the AI Recruiter that sources candidates from over 1 billion profiles, and writes about human-data vendors from the vantage point of someone whose product is often the alternative to renting them.
This guide reflects the human-data vendor landscape as of September 2026. Revenue figures, valuations, rates and leadership change frequently in this market; verify current details directly with each vendor before purchasing.








