We're hosting a researcher night next Wednesday at our office in SF. Fast-paced spotlight talks on post-training and evaluation. ⚡️ The talks: @mariyaivasileva on Beliefload: an evaluation of how LLMs formulate and revise hypotheses in light of new evidence, and how they diagnose and repair a broken environment without overwriting the parts that already work. @Nick_saban20 on GlobeBench: a benchmark for whether language models can faithfully simulate the environments we train agents in. @AnmolGulati06 on Beyond Rows to Reasoning: an agentic framework for reasoning over and editing enterprise spreadsheets with millions of cells, cross-sheet dependencies, and embedded charts. @akkikiki on SpeedrunBench: a benchmark that asks not whether an agent can finish a game, but how fast. Drinks, sushi, merch incoming.

Sep 17, 2026 · 6:17 PM UTC

1
6
21
2,142
Sort replies: Relevant Recent Liked