GM @SarvamAI | previously microsoft research, everwell, harvard university

For my last hurrah at Sarvam: super excited to launch Sarvam Vision 2.1! It's SOTA on olmOCR-Bench (87.3) and on our Indic OCR benchmark (87.39). With new capabilities like structured extraction and Indic handwriting recognition, it's a step change for document intelligence workflows 🧵
Sarvam Vision 2.1 is live, with new capabilities for structured extraction and Indic handwritten recognition across English and 22 Indian languages. The model is state of the art on both olmOCR-Bench and our own Indic benchmark, scoring 87.3 and 87.39 respectively. Read more in the blog: links.sarvam.io/zF7SiK
3
1
28
1,151
To support more rigorous benchmarking of Indian languages, we release the Sarvam Indic OCR Bench. It has 6,909 samples in 22 Indian languages and English, drawn from newspapers, textbooks, brochures and historical writing dated from 1800 to today. The benchmark contains samples with varying difficulty to measure OCR accuracy. Link: huggingface.co/datasets/sarv…
1
4
80
the talks from the vision track at AI Engineer WF 2026 are out! as part of an exciting line up showcasing cutting-edge work, i presented the Sarvam Vision 1.0 model
Live now: our Vision & OCR Track from AI Engineer World's Fair 2026. A model that counts 32 white squares on part of a chessboard. A file format that stores a table as a pile of line segments. Ten turkeys on the roof of a Tesla. Thesis: the models can see. They are still learning to look. piped.video/watch?v=RQi7x-na… - Building the Document Context Layer for AI Agents: @jerryjliu0, LlamaIndex - Skill issue: stop deploying vision language models, use them with Skills: @mervenoyann, Hugging Face - Modality Misalignment and Originality Attribution in Short-Form Video: Aditya Gautam, Meta - From Ingestion to Agents: How AI Teams Build on Document Intelligence: Adit Abraham, Reducto - The Best Models Still Reason Like Toddlers: @andrewdai, Elorian - You're Not Thinking Big Enough: Rebuilding Food Systems with AI Agents: @cbmenefee, Firecrawl - From VLM/VLA's to Embodied Agents: @ArmenAgha, Perceptron AI - From Scratch to SOTA: Training a 3B State-Space Vision Model: @fewshotlearner, Sarvam
1
4
255
football with constraints ❤️
1
96
Excited to present our work on Sarvam Vision at @aiDotEngineer World's Fair in SF. Over the last year, we have built India's first SOTA sovereign VLM for document intelligence. In my talk, I discuss what we did and how we did it. See you at Room 2006, Moscone Center. Happy to chat more about all the other work we are doing after the talk! @SarvamAI @SarvamForDevs @pratykumar
1
6
30
1,984
krishna retweeted
Sarvam joined the AI discussions at the 52nd G7 Summit in Évian-les-Bains, France. @vivekrag participated alongside G7 leaders and chief executives from some of the world’s leading AI companies. The discussion centered on frontier AI risks, infrastructure, and sovereignty in the age of AI. A significant moment for Sarvam and for India’s AI ecosystem.
46
281
2,421
69,309
krishna retweeted
We're thrilled to announce that we have raised $234M in the first close of our $300M Series B at a $1.5B valuation. @HCLTech and @BessemerVP have joined us in this round, alongside continued support from @khoslaventures and @peakxvpartners For countries and companies, sovereign control on the AI stack is no longer an optionality. Sarvam will be the partner of choice for this aspiration. The capital allows us to accelerate our momentum towards this full stack of models, compute, and deployments. A huge thank you to our customers, partners, investors, and the Sarvam team for your trust and belief in what we are building. We’re just getting started. Read more: sarvam.ai/announcing-series-…
648
1,496
9,904
1,108,403
Unlocking document intelligence for India scale efficiently!
Earlier this February, we launched Sarvam Vision, a vision-language model for document intelligence. Today, more than 35 million pages are being digitised through the Sarvam Vision API by developers and partners. Since launch, we've made it significantly more efficient to serve at scale. We’re now passing these gains on by reducing the Sarvam Vision API price from ₹1.5 to ₹0.5 per page.
2
3
234
Sarvam Vision, our SOTA document intelligence model, is now 66% cheaper! since launch, thousands of developers have digitized millions of complex documents with it. we've optimized our serving stack to scale with that demand - and we're passing on all of the gains to users. excited to see all that you build with the model. @SarvamAI @SarvamForDevs
Earlier this February, we launched Sarvam Vision, a vision-language model for document intelligence. Today, more than 35 million pages are being digitised through the Sarvam Vision API by developers and partners. Since launch, we've made it significantly more efficient to serve at scale. We’re now passing these gains on by reducing the Sarvam Vision API price from ₹1.5 to ₹0.5 per page.
1
13
867
eagerly waiting for series z and agi - whichever comes first 😂
We've raised $65 billion in Series H funding at a $965 billion post-money valuation, led by @AltimeterCap, Dragoneer, @Greenoaks, and @sequoia. This investment will help us advance our research and expand our capacity to meet growing demand for Claude.
1
126
the world of computer use agents (CUA) is gaining popularity. most frontier labs now have CUAs in some shape or form. SOTA on CUA is established using leaderboards - like OSWorld, AndroidWorld, etc. these include a battery of tests around click, tap, type, scrolll. things seemed good until i read this paper from Meta SI Labs. the team ran an experiment on the leaderboard - almost a prank! they took a CUA and let it solve a task; just once. they recorded every action into a 1mb file and showed a simple automation outperforms frontier CUAs. how so? performance in this domain essentially comes down to the test env. if every test begins from the same intial condition, then there is no perception needed. you dont need to see anything, or reason. mathematically this makes sense too: in a deterministic world, a simple automation or a complex agent system both yield the same result. the paper proposes a solution: randomize everything that can be randomized. each test should be executed in a fresh environment - a new sandboxed phone, new data, theme, UI state, so there is no gamification happening. read more: arxiv.org/abs/2605.08261
118
stay tuned for some really exciting updates to Sarvam document intelligence stack 🔥
Had a lot of fun attending @pratykumar's talk this week at Stanford. This is difficult engineering, done right at @SarvamAI! Also, got to know they will expand presence in the Bay Area-- All the best! :) @MohapatraHemant
1
1
217
these are getting ridiculously good
Anthropic onboarding day: Michael Scott introducing Karpathy like he just signed Wemby in free agency.
2
241
agree broadly with the thesis, but it's incomplete. data scarcity is only part of the problem. llms didn't learn to reason because the internet wrote reasoning down. chain-of-thought on the web before 2023 was minimal. they learned when post-training started using verifiable rewards (a math grader, a code runner, a unit test) to score intermediate steps. vlms have a similar gap, imo. there is no visual analog of the code runner. no verifiable check that asks "did the model actually see X in an image, or did it just say X based on the image context." if agentic vision were to take off, we should be build methods for verifiability. once a vision rl loop has a check for grounding, the diff vision skills become a normal post-training problem with rewards.
Since founding Moondream, I've watched language models achieve AGI, while VLMs aren't close to human-level visual reasoning. Here's why. 🧵
1
2
139
juxtapose elon's hiring call with the mass layoffs from meta. fascinating how the ai world is balancing itself out 😬
How founders need to hire in 2026. Though I'm curious how Elon is going to grok through the 10m applications while maintaining context to stack rank 😂
167