Almanack.
← Stories

Precedent · Differentiate & Position

Sarvam AI

Rule

Target the ground frontier players under-serve because it is technically awkward and politically defensible, then let a small team with the right pedigree own it before larger rivals bother moving.

Sarvam AI

Sarvam AI is an Indian startup founded in August 2023 by Vivek Raghavan and Pratyush Kumar to build "sovereign" AI for India: efficient, voice-first large language models that speak the country's languages rather than treating them as an afterthought. The through-line across its stories is a single strategic bet, that the ground the frontier labs under-serve (Indic languages, voice, and the trust of the Indian state) is both linguistically and politically defensible, and that a small team with the right pedigree could own it before anyone else moved in.

Beachhead: owning Indic voice, where frontier models tokenize badly

The problem. Global models such as GPT and Llama were built English-first, and they handle Indian languages inefficiently: they chop Hindi, Tamil, or Telugu into far more tokens per word than English, which makes them slower, costlier, and worse for the roughly billion Indians who do not transact comfortably in English. That left a large, unserved market that the frontier labs had little incentive to prioritize.

The approach. Rather than compete head-on on general English intelligence, Sarvam narrowed its wedge to voice-first agentic AI in Indian languages, purpose-building small, token-efficient models. Its Sarvam-1 (2 billion parameters) was trained from the ground up on 4 trillion tokens across ten Indian languages plus English, with a custom tokenizer designed to attack exactly that inefficiency.

How it solved it. Sarvam-1's tokenizer cut fertility to 1.4 to 2.1 tokens per word, close to the 1.4 typical for English, enabling 4 to 6x faster inference than larger models; on TriviaQA across Indic languages it scored 86.11%, far above Llama-3.1 8B's 61.47%. That efficiency turned into real business: Sarvam's voice and chat agents now automate KYC, sales, and support in local languages over telephony and WhatsApp for enterprises including Tata Capital and Infosys.

Why Now (Timing): catching India's sovereign-AI wave for compute it could not otherwise afford

The problem. Training foundational models from scratch is a capital and compute problem, and a young Indian startup could not easily match the GPU budgets of US labs. Sarvam needed thousands of scarce high-end GPUs at a moment when access to them was the binding constraint on frontier work.

The approach. Sarvam positioned itself squarely inside the Indian government's IndiaAI Mission and its push for a "sovereign" indigenous model, applying to build India's own LLM with state-subsidized compute. On 26 April 2025, MeitY selected Sarvam as the first startup, out of 67 shortlisted companies, to build India's sovereign foundational model.

How it solved it. The selection came with the largest compute subsidy in the program: roughly Rs 98.68 crore against a total bill of Rs 246.71 crore, granting access to 4,096 NVIDIA H100 GPUs for six months, plus a mandate to build a large open-source model for the country. Riding that same timing, Sarvam raised a $234 million Series B on 15 June 2026 (HCLTech alone put in $150 million for a 10.46% stake) at a $1.5 billion valuation, becoming a unicorn.

Founder-Market Fit: the Aadhaar architect and the AI4Bharat researcher

The problem. Building population-scale, government-trusted Indic AI demands two rare things at once: deep research in Indian-language NLP, and the credibility to operate at national scale with the Indian state. Very few teams could plausibly claim both, and without both, neither the government mandate nor serious capital would follow.

The approach. Sarvam's two founders embodied exactly that pairing. Vivek Raghavan had served as chief product officer and biometric architect on Aadhaar at UIDAI, India's billion-person digital-identity system; Pratyush Kumar was a longtime researcher (PhD from ETH Zurich, IBM Research), and together they had co-founded AI4Bharat, the open Indic-language AI lab at IIT Madras, before spinning out Sarvam.

How it solved it. That pedigree, national-infrastructure credibility plus Indic-NLP research depth, let Sarvam raise roughly $41 million in combined seed and Series A from Lightspeed, Peak XV, and Khosla Ventures in December 2023, within months of founding, and it underpinned the government's decision to entrust the sovereign-LLM mandate to Sarvam over larger incumbents. Its early OpenHathi work, adapting Llama and Mistral to Hindi, drew directly on the AI4Bharat lineage.