Sarvam AI releases Vision 2.1, claims top accuracy on Indian-language documents
SecondWing Editorial
4 min read

Bengaluru-based Sarvam AI has released Sarvam Vision 2.1, an update to its document-reading AI model that the company says now outperforms tools from Google, OpenAI, Anthropic and Mistral on Indian-language text. The release, announced on 24 September 2026, adds recognition of handwriting in Indian scripts and structured extraction from forms.
The model reads documents in English and all 22 scheduled Indian languages. It targets a problem that remains common across the country: records in regional scripts, often handwritten or poorly scanned, that most global OCR tools read unreliably.
What's new in version 2.1
The original Sarvam Vision launched on 5 February 2026, built on a 3-billion-parameter state-space vision-language model, according to the company and Business Today. Sarvam offered its document APIs free for that month to drive adoption.
In its announcement, Sarvam said feedback since launch fell into two areas: capabilities the model lacked, and practical problems running it in production. Version 2.1 makes four changes.
- Handwriting in Indian scripts. The company said it built training data from synthetic handwritten forms and from videos and other sources containing real handwriting.
- Key-value extraction from forms, returning fields such as name and date of birth rather than unstructured text.
- Better parsing of complex tables, including tables that run across several pages.
- Lower cost, fewer hallucinations. Sarvam said inference optimisations let it price the model below its launch rate, and that it worked to reduce invented or inconsistent output.
The underlying design is unchanged. A layout parser first identifies regions on a page, and a reading-order network sets their sequence before the model transcribes them. Sarvam described this approach as a deliberate trade in favour of accuracy.
How it compares
On Sarvam's own benchmark for Indian languages, Vision 2.1 posted the highest overall accuracy of the 13 systems tested. The chart below sets each model's English score against its Indian-language score, using figures Sarvam published.

By Sarvam's figures, frontier models including Opus 5 and Chat GPT 6 Astra lose 16 to 18 points when the text switches to Indian scripts, while Amazon's Textract scores under 5. Bodhan Indic-OCR, another Indian model, placed second. On the English benchmark the margin is narrow: Sarvam's 87.3 is about a point ahead of Infinity-Parser2 Pro.
Where it could be used
The new capabilities are aimed at sectors that still process large volumes of paper in regional languages.
- Banking and insurance: KYC documents, claim forms and handwritten applications, where form extraction applies directly.
- Government records and archives: Sarvam's benchmark draws on material dating from 1800 to the present.
- Accounting and small-business software that turns invoices and ledgers into structured data.
- Developers building for Indian-language users, who can call the model through an API rather than train their own.
Separately, at its Epoch 2026 conference, Sarvam made Vision Edge commercially available, a version that processes documents on local systems instead of a remote cloud, CXO Digital Pulse reported. That option may matter to hospitals and banks handling sensitive records.
Caveats
Several limits apply to the results, some of them acknowledged by Sarvam.
The results are self-reported. Sarvam ran every evaluation on its own setup, and no independent group has reproduced the scores yet. The Indian-language benchmark is also Sarvam's own, though the company has published it on Hugging Face for others to test against.
It is not a clean sweep. PaddleOCR-VL 1.6 scores higher on OmniDocBench. Bodhan Indic-OCR leads on Santhali by about 14 points, and Gemini 3.6 Flash is marginally ahead on Odia. Sarvam's own accuracy on Santhali and Kashmiri stays near 54%.
Scores across versions are not comparable. The Indic benchmark was cut from 20,267 samples at launch to 6,909, and the English results moved from a filtered subset to the official olmOCR set.
Benchmarks are near their ceiling. Sarvam itself noted that the main global benchmarks are close to saturated, so performance on real-world documents remains the more meaningful test.
Availability
Vision 2.1 is available now through Sarvam's API platform and a no-code playground.
Resource
Purpose
Document Intelligence Playground
Upload and test a document without code
Multi-page documents to structured text
Key-value pairs, tables and form fields
Human-in-the-loop document workflows
Open benchmark across 22 languages
Current rates
