Bodhan AI Releases Four Indic Models for OCR, Translation and Speech
AI Summary: Bodhan AI and AI4Bharat released four new models for processing Indian languages, covering document parsing, translation, speech recognition, and speech generation. The models, released in September 2026, include IndicOCR for parsing printed documents and handwritten text, Indic-Translate for document-level translation, Indic-Transcribe for speech-to-transcript, and Indic-Speak for text-to-speech. IndicOCR achieved 92.76 on OmniDocBench v1.6 and 86.2% word-level accuracy across 22 Indian languages and English, while Indic-Translate scored 58.97 dBLEU and 0.4326 word error rate on in-house document tests. The models support mixed languages and scripts, with applications in digitizing textbooks, making regional archives searchable, and localizing documents.