I will present this paper as a poster at COLM 2026 (Grand Ballroom #133, Franciscan, Poster #1199, F-4-133)!
HMU for research/coffee chats or with suggestions for good coffee shops! I've been working on making LLMs learn from non-language data, and I'll have more to say soon.
Language models that think, chat better.
We used longCoT (w/ reward model) for RLHF instead of math, and it just works. Llama-3.1-8B-Instruct + 14K ex beats GPT-4o (!) on chat & creative writing, & even Claude-3.7-Sonnet (thinking) on AlpacaEval2 and WildBench!
Read on. 🧵
1/8