Assistant Professor at IT University of Copenhagen

Denmark
Happy to share our paper, "Correct Answers, Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces". When a model gets the right answer, is the reasoning it produces also right? We test this on synthetic grade-school math where every step of a trace can be checked by a program. Thread below.
4
6
29
6,614
These findings matter for chain-of-thought monitoring in AI safety, where the trace is all a monitor sees.
1
1
4
174
I audited FineWeb-Edu, the educational subset of FineWeb, for misinformation. I annotated 200K documents from the 100BT sample using Llama 4 Maverick.
1
1
81
- On known problematic domains (naturalnews.com, rt.com, codoh.com, etc.), 39% of documents contain misinformation, nearly 10x the overall rate - Extrapolating to the full 1.53B document dataset: ~63M pages may contain misinformation
1
55
Ratish Puduppully retweeted
Getting scooped on my own paper drop by @DamienTeney 😂 But he's spot on - tabular language models have been overblown Reported performance gains often don’t come from emergent generalization at all. arxiv.org/abs/2602.04031
2
1
2
386
Happy to share our new paper! We study Tabular Language Models (TLMs): LLMs fine-tuned on serialized tabular data, converting rows into text sequences. Do TLMs actually learn tabular reasoning? Our re-evaluation suggests maybe not. 🧵 Paper: arxiv.org/abs/2602.04031
1
3
141
(3) Instruction-tuning alone with Alpaca (no tabular data) bridges most of the performance gap This suggests reported gains come from instruction-following and memorization rather than genuine tabular reasoning.
1
2
59
Great working with @ag0rla on this!
1
52
Excited to share our latest work on initializing expanded vocabulary for language models! Kudos to Nandini and Aditya for their excellent contributions!
🚨🚨 New preprint 🚨🚨 Presenting: An Empirical Comparison of Vocabulary Expansion and Initialization Approaches for Language Models Paper: arxiv.org/abs/2407.05841 Code: github.com/AI4Bharat/VocabAd… @anoopk @ratishsp @nandi_mundra
1
10
1,183
Ratish Puduppully retweeted
📣 📣 📣 New instruction-tuned LLM! 📣 📣 📣 Today, we announce an initial release of "Airavata", an instruction-tuned LLM for Hindi. Blog: ai4bharat.github.io/airavata… Model: huggingface.co/ai4bharat/Air… Datasets: huggingface.co/datasets/ai4b… (1/N)
8
83
362
51,633
Absolutely elated to be a part of this huge effort!
I am extremely pleased to announce that IndicTrans2 will be published in TMLR (@TmlrOrg). This is a tremendous achievement for my coauthors and me that took nearly 1.5 years of hard work. The camera ready version will be out soon but for now we are over the moon! #NLProc #ACL
7
675
Super happy to be a part of this work by @NameIsAshwanth!
Multiple acceptances in @emnlpmeeting thanks to students, researchers and collaborators. 1. CTQScorer: Combining Multiple Features for In-context Example Selection for Machine Translation Authors: @NameIsAshwanth, @ratishsp, @prajdabre1, @anoopk (Findings) #EMNLP2023 #NLProc
3
367
Thrilled to share that our paper on Decomposed Prompting for Machine Translation between Related Languages using Large Language Models is accepted to #EMNLP23 main conference! 🎉
1
1
15
1,350
Notably, DecoMT's performance improves as sentence length increases
1
2
140
Work done in collaboration with @anoopk, @prajdabre1, Ai Ti, and Nancy.
2
124