Ask a frontier AI model a question in English and it answers like a well-read consultant. Ask the same model in Maithili — a language with more native speakers than Greek — and it stumbles like a tourist with a phrasebook. For the sixty million people who think in Maithili, the most important technology of the century arrived speaking someone else’s language.
That gap is now the most contested territory in Indian technology.
The data problem underneath everything
Large models learn from text, and the internet’s text is catastrophically lopsided. English accounts for roughly half the web’s content; all twenty-two scheduled Indian languages together account for well under one percent — despite representing a billion speakers. Hindi, spoken by more people than any language in Europe, has less online text than Danish.
The response has been a national data harvest unlike anything attempted elsewhere. Government-backed missions are recording tens of thousands of hours of speech in dialect belts where “Hindi” means six different things. Publishers are digitising a century of regional literature. AII-IT crawls now treat a Bhojpuri folk-song transcript as strategic infrastructure, because that is precisely what it has become.
A language without data is a language without a future tense. The harvest is a hedge against silence.
Small models, sharp targets
The teams building Indic models have largely stopped chasing the frontier labs at their own game. The winning pattern is smaller models, trained dense on curated Indic corpora, tuned for the tasks that actually matter at the panchayat level: a farmer querying crop insurance by voice, a patient describing symptoms to a telehealth line, a student getting maths explained in the language her mother uses.
Voice is the battleground, not text. For hundreds of millions of Indians, the first interface with AI will be spoken — which means the models must handle code-switching, dialect drift, and sentences that begin in Kannada and end in English, because that is how India actually talks.
Who gets to be understood
The stakes are easy to understate as a product problem. Whoever builds the models that speak India’s languages fluently will mediate credit applications, government schemes, medical advice, and the homework of a generation. If those models are built elsewhere, tuned on translation rather than lived usage, a billion people get a machine that perpetually almost-understands them.
The engineers in Bengaluru’s Indic-AI labs put it more simply. English got the first draft of this technology. The other twenty-two tongues are owed the final one.


