AIKosh: Scriptures and Parliament Data for Indian AI
Why in the news
The government intends to feed religious and cultural texts in many languages, plus public governance records, into AIKosh so that Indian AI and large language models learn from local context.
Key facts
- Languages: Hindi, English, Tamil, Telugu, Kannada, Urdu and more; also local dialects and oral storytelling.
- Rationale: scriptures hold ancient wisdom that can improve accuracy, cultural alignment and contextual depth.
- MeitY signed an MoU with the Lok Sabha Secretariat for Parliament questions and answers, government reports, committee papers and ministry agendas.
- Other inputs: non-personal, anonymised datasets and the Open Governance Data Platform.
| Item | Detail |
|---|---|
| Datasets (April 9) | 350+ |
| AI models hosted | Nearly 150, both LLMs and SLMs |
| Parent programme | India AI Mission, ₹10,372 crore |
| Budget 2025-26 allocation | ₹200 crore |
About AIKosh
- It is the India Datasets Platform, one of seven core pillars of the India AI Mission.
- No monetisation of datasets by government or private players, as MoS Jitin Prasada told Parliament; no data purchases or subscriptions.
- Follows strict data protection standards and Indian laws, including the IT Act, 2000 and the Data Protection Bill.
Exam angle
- Nodal ministry: MeitY.
- MoU partner: Lok Sabha Secretariat.
- Platform is part of the India AI Mission’s seven pillars.