Applied AI @SnorkelAI, Previously - @MistralAI , @llama_index (LlamaIndex) Building data and evals for frontier labs.

SF
AutoResearch experiment: I went into a GPU Mode kernel competition with very little CUDA background and engineered a research loop around coding agents: goal, rules, experiment memory, verifier, and compute. Finished 5th overall. Give it a read.
Article

AutoResearch: 5th place in GPU Mode with coding agents

I recently finished 5th in a GPU Mode competition for batched Cholesky factorization on B200. The final submission scored 265.234 μs, down from a baseline of 2533.030 μs, after 500+ official

30
36
354
44,084
Ravi Theja retweeted
🥳 Excited to share that @lossfunk got 4 papers into NeurIPS main conference. Please join me in congratulating the authors! Aman @arcaman07, ⁠Sushrut @martisamuser, ⁠Abhinav @MajorTimbWlf21, Pranav - @Avg_sapient First authors are all undergrads :)
59
70
1,485
73,608
Ravi Theja retweeted
I'm excited to announce @SnorkelAI's $350M Series E at $3.5B, led by @insightpartners and @S32_VC. We've grown 18x+ in the last 12 months since launching our Data-as-a-Service offering, passing $375M ARRR this week. As AI advances to superhuman capabilities, AI data & environment development must advance with it - and basic staffing and crowdsourcing approaches are not enough. AI progress now requires deep research and technology work that combines human expertise with specialized AI in compounding ways. @SnorkelAI is building the RSI data engine and frontier data lab for this next phase. We're honored to have the support of existing investors Addition, @lightspeedvp, @GreylockVC, @GVteam, P7, Factory, @WellsFargo, Walden Catalyst Ventures, and new investors @ThirdPointLLC, @MarchCPs, @BlumbergCapital, @AllegisCapital, @Frontlinevc, and @standard_vc. – @SnorkelAI started as a research project a decade ago at @StanfordAILab. Our thesis was simple: AI progress would become increasingly data-centric – and therefore data development should be studied as a true research and technology problem, not just a staffing and crowdsourcing one. Today, as AI capabilities verge on superhuman, building the data and environments to safely measure and train AI is becoming too hard for even the smartest human experts to do alone. Only humans and AI agents, collaborating together in compounding ways, can meet the accelerating needs of the frontier, and keep humans in the driver’s seat of AI progress for decades to come. At @SnorkelAI, we are building the data lab to define the shape of this new “Data 2.0” frontier, and the new paradigms of human-computer interaction needed to advance it. Our key focus is building the RSI engine for data, where specialized AI models accelerate and improve human expert output, and in turn, scaled human supervision is used to continuously evaluate and improve these models – creating a powerful compounding loop to keep pace with an accelerating RSI frontier. With this round of funding, we are also doubling down on our commitments to support data development for open benchmarking and evaluation (more news here soon!); an increasingly diverse ecosystem of general and specialized intelligence; and a path to safe, well-aligned AI built on robust training and evaluation data. Data development will guide and drive the next stages of AI – and must do so in a human-centric, AI accelerated, open, diverse, and safe way. We are excited to support this mission in the next decade of research ahead at @SnorkelAI. More thoughts here: snorkel.ai/blog/data-2-0-and…
72
115
410
88,054
We thought AI needed good computer-use models to play games. Turns out, all it needed was Jev.
jev is insane 🤯 Here is Jev playing subway surfers at super human speed, and also playing 50 games at once. cost less than a cent to do this run. Jev does not replace llms like astra or fable, but opens up an entirely new world of capabilities.
1
1
14
2,004
Interesting that 9/15 reviewed benchmarks are reported as "Flawed"
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
6
792
Ravi Theja retweeted
i'm flattered by the shoutout, but there's no free lunch for document parsing Cohere parse starts at $1.50 / 1k pages, which is indeed on the cheaper end of doc OCR solutions, but it scores ~50% when averaged across all 5 ParseBench dimensions: * It lacks high-quality visual grounding, doing coarse segmentation for tables and images, but lacking finer-grained bboxes at the word/cell/line-level. this is quite important for any sensitive AI application that requires precise citations back to the source. * It lacks general visual / chart parsing capabilities, which are prevalent across finance, manufacturing, healthcare, and more * to its credit, it does a decent job over tables (87%) compared to other comparable solutions at the price point our cost-effective mode in LlamaParse starts at 0.375c / page and is more consistently strong across all dimensions. also there's various ways to bring it down with volume discounts if you're interested in learning more about various document parsing modes and comparing tradeoffs in price/accuracy, come talk to us! llamaindex.ai/contact LlamaParse: cloud.llamaindex.ai/
You're paying too much to parse your documents. Cohere Parse 5 fixes that.
22
11
121
25,799
Recurrence Is Not Enough: Causally Validating Multilingual SAE Translation Features in Gemma 2 and 3 Giang Son Nguyen, Nhi Ngoc-Yen Nguyen, Wray Buntine, Dung D. Le arxiv.org/abs/2609.04808 [𝚌𝚜.𝙲𝙻 𝚌𝚜.𝙰𝙸] 💬Accepted to BlackboxNLP 2026 Special Track
1
114
Building games with GPT-6 Astra is so much fun! Built a racing game set on Bengaluru’s ORR, from Bellandur to Marathahalli. Tried to capture the familiar chaos: traffic congestion, potholes (h/t @HarveenChadha), Swiggy, Zomato, and Blinkit delivery riders, and metro construction. And for some Road Rash nostalgia, you can swing a rod at vehicles alongside you to make room 😂 (h/t @ptkbhv). Give it a try: orr-rush-bengaluru.ravitheja…
4
4
23
1,430
Ravi Theja retweeted
You can just do things.
53
390
3,060
135,381
Ravi Theja retweeted
As engineering, product, design, DS, etc. melt into a new kind of role, I was reflecting on what roles might look like in the future. For example, when I look at the Claude Code team I see what I think is five archetypes: 1. Prototyper: comes up with brand new ideas; churns out many ideas, most of which don't ship 2. Builder: quickly turns a prototype/idea into production-grade product/infra 3. Sweeper: cleans up the UI, simplifies the code and system, unships, optimizes performance 4. Grower: takes a product that has been built and iterates on it to improve Product-Market Fit 5. Maintainer: owns a mature system to make it secure, reliable, fast, and efficient as it scales Many people span across 2 roles, and sometimes 3 roles. I also notice that these roles are not really tied to job function -- eg. across Anthropic, some designers match category 1, some 2, some 3; same for engineers, PM, DS. A healthy team needs a mix of these, depending on the product: - A product that is new and pre-PMF needs people that are strong at 1+2+3 - A product that is growing and has found PMF needs 2+3+4 and some 5 - A product that has strong PMF needs 3+4+5 and some 2 Maybe product roles of the future will look more like this, and less like the domain-specific roles of today?
886
2,442
20,053
3,150,364
Ravi Theja retweeted
We are releasing a fully reproducible early preprint of "Prism: Unlocking Language Model Capability Extraction". A trained language model knows many things at once, but deployment usually asks for one behavior at a time. Enterprise scenarios often have few products, workflows, features, or use-cases matter disproportionately. Prism asks and answers a simple question - "Is it possible to isolate and deploy only capabilities that are driven by Pareto principle and cut down costs by a huge margin while preserving most of the performance?" This paper discusses a novel approach to efficiency, understanding model behavior and opens up capability extraction.
21
38
221
24,740
Ravi Theja retweeted
We're presenting ParseBench at CVPR 2026! ParseBench is the most comprehensive document understanding benchmark for VLMs. ✅ It contains 2k pages of real-world enterprise documents ✅ It has comprehensive evaluation metrics around tables, charts, visual grounding, semantic formatting, and content faithfulness The core goal is measuring whether models can semantically interpret a document in the right way, without having models overfit to our precise benchmark. Parsing 100% of PDFs to 100% accuracy is the final boss for document OCR. In general, the latest frontier models have been tuned for coding, math, and scientific reasoning as opposed to precise visual understanding; hope more benchmarks that these will encourage overall progress towards solving this problem! Poster is below. If you want to learn more come check out our site or 30-page ArXiv paper: ParseBench: parsebench.ai/ ArXiv: arxiv.org/abs/2604.08538
We're presenting ParseBench at CVPR 2026 today. 🦙 Come learn why document understanding is an AGI-complete problem (an agent can't act on a doc it can't correctly read, and reading a real enterprise table is harder than it looks). The first doc-parsing benchmark built for AI agents: 2,000+ human-verified pages 167K+ test rules 5 dimensions: tables, charts, faithfulness, formatting, grounding Fully open source. 📍 Talk TODAY, June 4, 9–10 AM at CVPR. Come say hi 👇 🤗 huggingface.co/datasets/llam… 💻 github.com/run-llama/ParseBe… 📄 arxiv.org/abs/2604.08538
12
17
106
18,552
Ravi Theja retweeted
Superintelligence will be built on Self Improvement. Today @hexoai, we’re excited to release ‘SIA’ - an open-source Self-Improving AI, to achieve any goal through recursive self improvement. While trying to solve a problem, SIA doesn't just improve it's abilities by updating it's harness, it updates it's own weights as well.
120
132
635
529,379
In another decade, this might be a robot. For now, Indian unit economics favor the human version :)
In Lajpat, you can now pay ₹149/hr for someone to carry your bags, wait in food queues, walk you to the metro, find you a place to sit, and even set up a foldable chair. interesting biz !!
1
3
1,187
Ravi Theja retweeted
For a while, @aakrit & I've quietly done this at Activate: spend time with exceptional operators, CTOs & senior builders before there’s a deck, company, or fully formed idea. Now opening it up Mixture of Experts - a small private dinner for people at the edge of founding First edition tomorrow in Bengaluru Request below on the Luma to join or DM if someone exceptional should be in the room. luma.com/1vmm7hq9
6
3
93
18,806
Ravi Theja retweeted
Ever wondered if you could extract capabilities and behaviors from neural networks and reuse/update/route it as needed? We introduce low-rank circuit conditioning, a novel approach that preserves the model's output behavior while reshaping how an existing capability is represented. In the base model, standard compact recovery stalls at 29%. After conditioning, the same extraction pipeline reaches 91.33% autoregressive full-answer recovery from 5.05% of MLP channels. The evidence points to a possibility of extracting and using isolated capabilities saving cost, latency and high adaptability. Read our work to understand more - tokenbender.com/posts/honey-…
28
65
393
45,414
Ravi Theja retweeted
India’s AI future is being built by those who ship, not just study. Today I’m announcing Activate Fellows - a summer program for 15 of India’s (and the world’s) best student builders to work inside the country’s leading AI startups Only 15 spots. And 8 days left to apply. Details + link 🧵
10
18
183
15,160
Ravi Theja retweeted
ParseBench is now live on Kaggle Benchmarks! 🚀 Developed by @llama_index, this benchmark evaluates PDF-to-structured-data conversion, featuring ~2k human-verified pages from real enterprise docs across 5 capability dimensions. 🥇Gemini 3 Flash: 79.3% 🥈GPT 5.4: 72.9% 🥉Gemma 4 31B: 66.4%
5
19
116
15,654
Ravi Theja retweeted
We’re open sourcing the first document OCR benchmark for the agentic era, ParseBench. Document parsing is the foundation of every AI agent that works with real-world files. ParseBench is a benchmark that measures parsing quality specifically for agent knowledge work: ✅ It optimizes for semantic correctness (instead of exact similarity) ✅ It has the most comprehensive distribution of real-world enterprise documents It contains ~2,000 human-verified enterprise document pages with 167,000+ test rules across five dimensions that matter most: tables, charts, content faithfulness, semantic formatting, and visual grounding. We benchmarked 14 known document parsers on ParseBench, from frontier/OSS VLMs to specialized parsers to LlamaParse. Here are some of our findings: 💡 Increasing compute budget yields diminishing returns - Gemini/gpt-5-mini/haiku gain 3-5 points from minimal to high thinking, at 4x the cost. 💡 Charts are the most polarizing dimension for evaluation. Most specialized parsers score below 6%, while some VLM-based parsers do a bit better. 💡 VLMs are great at visual understanding but terrible at layout extraction. GPT-5-mini/haiku score below 10% on our visual grounding task, all specialized parsers do much better. 💡 No method crushes all 5 dimensions at once, but LlamaParse achieves the highest overall score at 84.9%, and is the leader in 4 out of the 5 dimensions. This is by far the deepest technical work that we’ve published as a company. I would encourage you to start with our blog and explore our links to Hugging Face to GitHub. All the details are in our full 35-page (!!) ArXiv whitepaper. 🌐: Blog: llamaindex.ai/blog/parsebenc… 📄 Paper: arxiv.org/abs/2604.08538?utm… 💻 Code: github.com/run-llama/ParseBe… 📊 Dataset: huggingface.co/datasets/llam… 🎥 YouTube: piped.video/watch?v=g5p7G-Nw…
33
81
521
108,333