Win one narrow market that desperately needs quality nobody else can deliver, then ride that same specialized pipeline into every successive wave as the picks-and-shovels supplier everyone depends on.
Scale AI
Scale AI is the data-labeling and model-evaluation infrastructure company founded in 2016 by 19-year-old MIT dropout Alexandr Wang and Lucy Guo, built on the conviction that data, not compute or algorithms, is the real bottleneck in machine learning. Its story is one of owning the unglamorous data layer: it wedged into a single desperate market (self-driving cars), then rode its labeling pipeline forward into every subsequent AI wave (generative models, RLHF, evals, defense), becoming the picks-and-shovels supplier whose economics culminated in Meta's 2025 stake.
Beachhead: labeling LiDAR for self-driving cars nobody else could handle
The problem. In 2016, autonomous-vehicle programs had huge compute budgets but no reliable way to turn raw sensor feeds into training data. Labeling 3D LiDAR point clouds and video frame by frame was intricate, high-stakes work that general-purpose crowdsourcing vendors could not deliver at the quality or scale self-driving required.
The approach. Rather than build a broad labeling platform, Scale aimed narrowly at AV companies, offering an API where a developer sent raw sensor data and got back precise annotations. Wang and his team went booth-to-booth at the 2016 Computer Vision and Pattern Recognition (CVPR) conference, demoing the platform on laptops to the exact researchers who needed it.
How it solved it. The wedge landed: early clients included Toyota Research Institute and Lyft, and Scale became a go-to annotation partner for Cruise, Nuro, Waymo, Uber, and General Motors. By choosing the one vertical most starved for labeled data, Scale built a defensible foothold that carried it to a $1 billion valuation and a $100 million Series C led by Founders Fund in 2019.
Land and Expand: from car sensors to the generative-AI data engine
The problem. The autonomous-vehicle market matured slowly and remained a limited pool of customers, while the frontier of AI shifted after 2022 toward large language models, which needed a completely different kind of human data: written responses, preference rankings, and reinforcement learning from human feedback (RLHF).
The approach. Scale reused its core asset, a managed pipeline of human labelers plus tooling, and pointed it at generative AI, launching a Generative AI Data Engine for RLHF, human data generation, model evaluation, safety, and alignment. It expanded the same land-and-expand motion into new accounts and adjacent verticals, including government and defense AI programs.
How it solved it. The expansion turned Scale into the data supplier behind leading labs including OpenAI, Meta, and Microsoft, exactly as those companies raced to fine-tune LLMs. The pivot is visible in the numbers: revenue reached roughly $870 million in 2024, with the company projecting over $2 billion for 2025, growth driven overwhelmingly by generative-AI work rather than the original AV business.
Business Model: selling shovels to every side of the AI gold rush
The problem. AI labs compete fiercely and their model architectures change constantly, so betting the company on any single model or lab would be fragile. Scale needed a business model that profited from the whole industry's growth regardless of which lab won.
The approach. Scale positioned itself as neutral, usage-based infrastructure: it charges for data services (labeling, RLHF, evaluations) that essentially every model builder depends on, making itself the picks-and-shovels layer beneath the AI boom rather than a competitor in the model race.
How it solved it. By becoming the shared data supplier to rivals across the field, Scale made itself strategically valuable enough that in 2025 Meta invested $14.3 billion for roughly a 49% stake, valuing the company at over $29 billion and pulling founder Alexandr Wang into Meta to lead its superintelligence efforts. That outcome validated the thesis Wang started with: own the data layer, and you own leverage over everyone building on top of it.
