LLM & language operations

The human-built data behind the IIT-Madras research lab building open-source Indian-language AI

The client is the IIT-Madras research lab building open-source AI for India's languages, from translation to speech. DesiCrew has been its data partner since 2019, building the transcription, translation and corpus data those open models are trained on, across 22 languages.

The human-built data behind the IIT-Madras research lab building open-source Indian-language AI
Since 2019
The lab's data partner
22
Scheduled Indian languages covered
100+
In-house linguists behind the data
1B+
Indic speakers the models are built for

The client is the IIT-Madras research lab building open-source AI for India's languages, from translation to speech. DesiCrew has been its data partner since 2019, building the transcription, translation and corpus data those open models are trained on, across 22 languages.

A model is only as fluent as the data behind it, and for most Indian languages that data did not exist. DesiCrew's linguists built it, by hand, language by language. So an open model can understand a farmer in Odia or a patient in Telugu, and anyone can build on it.

01 — The Mandate

The lab builds open-source AI for Indian languages, and open models need open, high-quality data underneath them. DesiCrew produces that data.

Scale. 36,000+ hours of transcription, 10,000+ hours of speech data collected, 900,000+ words translated, and 1.1 million+ words of LLM training data built. Coverage. 22+ Indian languages, 100+ in-house linguists, and a 1,500-strong workforce behind them. Lines of work. Custom speech and text data collection, translation and post-editing, time-aligned transcription and segmentation, and LLM support work like Q&A generation, RLHF and responsible-AI datasets. Grounded in research. Methods shaped by 30 to 35 research papers, because the data feeds an academic program that publishes its work and is examined on it. Open by default. Everything built to be released — the datasets and models are open-source and free, so the data has to be clean enough to publish, not just to train on.

02 — The Challenge

Recording the audio is the easy part. Making it data a model can actually learn from is the hard part. Most of the world's AI is trained in English and a handful of high-resource languages. India speaks hundreds, most with almost no machine-readable data behind them. Without it, a model cannot understand a farmer asking a question in Odia or a patient describing symptoms in Telugu. Building that data is slow, human work — and the model is only ever as good as the data underneath it.

The data does not exist yet. For a low-resource language there is no clean corpus to pull from; every hour of speech and every line of translation has to be collected, transcribed and verified from scratch. Language is not just words. A dataset has to carry dialect, context, the code-switching between English and the local language, and the meaning a literal translation would lose — that needs linguists, not typists. Bias and safety are set here. What a model learns about a language, and about the people who speak it, comes from this data; responsible-AI datasets, bias detection and context-aware data are part of the build, not an afterthought. It goes out in the open. The lab's data and models are released publicly and free — an error does not hide in a private dataset, it ships to every researcher, startup and government team that builds on the work. It has to survive scrutiny. The team has drawn on 30 to 35 research papers to shape its methods, because the approach has to be defensible, not just fast.

Anatomy of one language dataset — why a model cannot build it itself

One commissioned dataset can run from 1,000 to 8,000 hours — collected as speech, time-aligned to transcription, segmented, translated, post-edited, and checked for bias and context, across 22+ languages, by 100+ linguists who actually speak them. A model can generate text. It cannot decide what is true in a language it was never taught: whether a translation carries the right meaning, whether a dialect is represented, whether the data is safe and unbiased.

03 — The Approach

DesiCrew collects the data that does not exist, puts linguists where a model cannot judge, and builds it to the open standard the client publishes to.

Collect at source. Custom speech and text data collection in the target language, gathered fresh where no usable corpus exists. Transcribe and align. Time-aligned transcription and segmentation, so audio and text line up precisely enough to train the speech models. Translate with linguists. Translation and post-editing of machine output by 100+ in-house linguists across 22+ languages, so meaning survives, not just words. Build for the model. Q&A generation and text-corpus creation, shaping raw language into the parallel and monolingual data the models actually train on. Responsible by design. Bias detection, responsible-AI and context-aware datasets, built DPDPA-compliant and grounded in 30 to 35 research papers.

04 — The Outcome

Open models trained on real Indian-language data. The datasets DesiCrew built feed open translation, language and speech models rather than data scraped and hoped over — so Indian-language AI has a public foundation, not English translated badly.

22 languages, most of them low-resource. Data built across every scheduled Indian language, including those with almost no machine-readable text before the work began — so a farmer in Odia or a patient in Telugu is inside the model, not left out of it. A public dataset the whole country builds on. Because the data is open and free, the work compounds beyond one model — into startups, research and national language programs. So one dataset seeds an entire ecosystem of Indian-language AI.

05 — Why the relationship holds

DesiCrew adopted this partnership in 2019. The moat is the linguists. 100+ of them, in-house, across 22 languages, is not something a general annotation vendor can assemble on demand. That fluency is why the data holds up to the research scrutiny.

About DesiCrew

DesiCrew is an applied-intelligence company — the human and technology layer that makes AI systems and enterprise operations work reliably and grow at scale, refined in production since 2007. IIT Madras incubated · Everest Group PEAK Matrix 2024 · Great Place to Work.

Let's put intelligence to work.

If your AI keeps breaking when it leaves the lab, or your operations are carrying weight AI should be taking off — let's build together.