Snorkel AI has raised $350 million in a Series E round that values the training-data company at $3.5 billion, giving it fresh capital as frontier model developers demand more specialized datasets, simulated environments and evaluation work.

 

The financing brings three closely watched signals:

  • A valuation nearly three times Snorkel AI's May 2025 level.
  • Rapid growth in a data-as-a-service business launched last year.
  • Investor confidence that advanced models still need expert human input.

 

Snorkel AI Raises $350 Million in Series E Funding

Insight Partners and S32 led the round. Snorkel named Third Point, March Capital, Blumberg Capital, Allegis Capital, Standard VC and Frontline Ventures among the other participants, alongside existing backers including Addition, Lightspeed, Greylock, GV and Wells Fargo.

 

The company said the financing values it at $3.5 billion. Reuters reported that this is nearly triple the $1.3 billion valuation Snorkel reached when it raised $100 million in May 2025.

 

Snorkel also said its annualized revenue run rate crossed $375 million during the week of the announcement, after growing more than eighteenfold since it launched its data-as-a-service offering nearly a year earlier. That figure is company-reported rather than audited annual revenue.

 

The distinction matters because run rate extrapolates recent sales performance over a full year. Even so, the figure indicates how quickly demand for specialized AI training and evaluation material has expanded beyond traditional data-labeling work.

 

How Snorkel's Agentic Data Platform Works

Snorkel began as a Stanford research project focused on weak supervision, a method for using rules and other imperfect signals to label data at scale. The company later commercialized software for enterprise machine-learning teams before expanding into finished datasets and reinforcement-learning environments.

 

Its current platform combines subject-matter experts with specialized models and agents. Experts define scenarios, tasks, grading criteria and failure conditions, while automated systems help generate candidate material, route reviews and perform repetitive quality-control checks.

 

That hybrid model is aimed at a growing bottleneck. Stronger models can solve many routine tasks, so useful training examples increasingly need to target narrow weaknesses, long workflows and difficult edge cases that ordinary web data does not capture.

 

Coding is one of Snorkel's largest demand areas, according to Reuters. The company also draws on specialists in law, medicine and other technical fields, where evaluation can depend on professional judgment rather than a simple correct-or-incorrect label.

 

Why Frontier Models Need Harder Training Data

Pretraining large models depended heavily on collecting vast quantities of text, code and images. Post-training newer systems increasingly requires curated examples, interactive environments and reward signals that teach a model how to complete multi-step work safely and reliably.

 

For coding agents, that can mean reproducing a realistic software repository, specifying a task that takes hours to solve, identifying acceptable outcomes and detecting shortcuts that satisfy an automated test without completing the intended work.

 

Snorkel says hundreds of specialized agents can support quality control for some task types, while human reviewers remain responsible for defining goals and judging ambiguous outputs. The company argues that neither purely manual production nor purely synthetic generation can meet frontier developers' needs alone.

 

This market has drawn major investment because data quality can limit model improvement even when computing capacity and model architecture continue advancing. Scale AI, Mercor and Surge AI are among the companies competing to supply labs with expert feedback and complex training tasks.

 

Where Snorkel Plans to Spend the New Capital

Snorkel plans to hire more researchers and engineers, expand its enterprise and U.S. government operations, support third-party model evaluations and move into additional industries and data types. The company works with frontier labs, cloud providers, specialist AI companies and public-sector customers.

 

Its next challenge is turning rapid project-based demand into a durable platform business. Customers building frontier systems may need highly customized environments, while enterprise buyers typically expect repeatable products, security controls and predictable delivery.

 

Snorkel told Reuters that it expects to reach profitability in 2026, even as growth remains the priority. That target will test whether its mix of expert labor and automation can preserve margins as tasks become more technically demanding.

 

The round does not prove that one data-production method will dominate. It does show that investors see training data, evaluations and reinforcement-learning environments as a separate infrastructure layer with enough demand to support multibillion-dollar companies.

 

Related Research