We're hosting a researcher night next Wednesday at our office in SF. Fast-paced spotlight talks on post-training and evaluation. ⚡️
The talks:
@mariyaivasileva on Beliefload: an evaluation of how LLMs formulate and revise hypotheses in light of new evidence, and how they diagnose and repair a broken environment without overwriting the parts that already work.
@Nick_saban20 on GlobeBench: a benchmark for whether language models can faithfully simulate the environments we train agents in.
@AnmolGulati06 on Beyond Rows to Reasoning: an agentic framework for reasoning over and editing enterprise spreadsheets with millions of cells, cross-sheet dependencies, and embedded charts.
@akkikiki on SpeedrunBench: a benchmark that asks not whether an agent can finish a game, but how fast.
Drinks, sushi, merch incoming.
Sep 17, 2026 · 6:17 PM UTC
1
6
21
2,142

