Applied research lab curating data solutions to accelerate foundation model development.

Based in United States
Congratulations to the @motif_tech team on the release of Motif 3! AfterQuery is proud to have served as the sole data partner on this model.
9
6
40
14,313
Congrats to the @WeAreLegora team on the release of the Legora BAR, a benchmark for agentic legal work built on 5,161 real law firm cases across 28 practice areas! @AfterQuery helped QA the benchmark with Legora. Proud to work with their team on measuring the frontier of agentic legal work. Link to the full blog post in the replies.
11
11
35
28,361
Last Thursday, we hosted Poker Night at our SF office. Researchers, founders, and builders joined the @AfterQuery team for 3 things: 1. Poker. 2. Drinks. 3. Lively conversation. We're hosting more events soon. Comment if you want to join the next one ↓
15
8
40
8,129
See what your codebase could be worth. Introducing Atrium by @AfterQuery. We’re paying companies for their private codebases. Connect a repo, get an automated quality analysis in minutes, and receive a quote in days. Link in the replies.
15
14
117
32,336
Kimi K3 ranks #1 on @AfterQuery's SpreadsheetBench 2, surpassing Claude Fable 5. An open weight model now outperforms all closed-sourced models. Read more in the Kimi K3 blog and SpreadsheetBench 2 paper linked below. Congrats to the @kimi_moonshot team on the incredible model!
37
119
1,044
147,549
The @AfterQuery team is in Korea this week for ICML, but the best conversations didn’t happen in the conference halls. They happened around a chef’s table. This week, we partnered with @altosvc to bring together researchers, founders, and investors for an evening of great food, drinks, and meaningful conversation. More than 100 people joined us for the evening, with a special menu prepared by Culinary Class Wars Lim hee-won and cocktails from Zest, named Asia’s Best Bar in 2025. Huge thank you to @altosvc for partnering with us, and to all the attendees who made the event special. We look forward to hosting more events in the future that bring the AI research community together.
4
4
19
2,983
Introducing @LM_Arcade - play, rank, and submit AI generated games. Supported by @AfterQuery. Models generate games using several popular agentic harnesses. Since tools, agent loops, and scaffolding can meaningfully affect performance, we evaluate both the models and the harnesses rather than forcing every model into the same environment. Current top 3 models: @AnthropicAI Claude Fable 5, @OpenAI GPT-5.5, @AnthropicAI Claude Opus 4.7 Top 3 harnesses: @AnthropicAI claude-code, @AfterQuery aq-gaming, @Princeton mini-swe-agent. Blind voting: compare two game builds from the same prompt without knowing which model or harness created them. Vote for the one that’s more enjoyable and better matches the request. Bradley-Terry ratings: every vote updates the leaderboard in real time. Don't just read about performance. Determine it at lmarcade.com.
5
2
18
7,294
Congratulations to @nvidia on the release of Nemotron 3 Ultra! Ultra now leads GDPval-AA by a good margin among US open source models. @AfterQuery is proud to have provided the GDPval training data used to hill climb the benchmark.
9
9
35
11,951
As corporate "tokenmaxxing" efforts struggle to generate meaningful enterprise value and model capabilities begin to commoditize at the frontier with cheaper open source options quickly approaching in the rearview mirror, @OpenAI and @AnthropicAI are scrambling to assemble FDE armies with strategic PE partners to solve the last mile enterprise problem and sustain their exponential revenue growth. Read our full take on OpenAI’s DeployCo and Anthropic’s ServiceCo at afterquery.com/blog/deployco
1
5
15
2,283
YC and @GoogleDeepMind are hosting the Multimodal Frontier Hackathon this Saturday. Most AI apps still don't utilize the full multimodal stack. So we’re giving you access to Gemini 3.1, Lyria, & NanoBanana 2 to see what you can build! Sign up at: events.ycombinator.com/deepm…
2
11
3,316
Introducing IDE-Bench! A multi-language, full-stack benchmark evaluating LLMs acting as autonomous IDE agents IDE-Bench assesses agents' ability to navigate, reason, and modify complex repositories using the same tools available in modern AI-native IDEs like Cursor Models tested from @AnthropicAI, @OpenAI, @Alibaba_Qwen, @GoogleDeepMind, @xai, @deepseek_ai, @Meta, and @cohere Check out the full results at ide-bench.com!
8
9
21
2,369
Introducing Market-Bench by @AfterQuery! The first-of-its-kind benchmark on LLMs for quantitative finance. We challenged models to attempt a frequent introductory quantitative trading task: coding an executable backtester from a natural-language strategy description and market assumptions. > 13 models build backtesting systems for directional, pair trading, and delta hedging strategies > evaluated on reliability (executable passes) and accuracy (MAE) across 5 attempts per strategy > real order book data with exchange delays and liquidity constraints > @xAI’s Grok 4 achieved the overall lowest mean MAE (deviation from the golden backtest), followed closely by @OpenAI’s GPT 5.2 > @AnthropicAI's Sonnet 4.5 and @AlibabaGroup's Qwen 3 Max at perfect executability but high MAE > Models from @Meta, @Amazon, @NVIDIA, and @Cohere continued to fail to produce executable backtesters Leaderboard & full paper below!
5
2
13
1,781
Today, humanity is shackled by scarcity of expertise. When expertise becomes infinitely scalable, humans will be freed to tackle problems we can't even conceive of today. Introducing @AfterQuery. We’re building a world where expertise is abundant. Domain by domain, profession by profession, AfterQuery is crafting datasets that encode excellence into forms that machines can learn. Data is the final frontier.
77
16
72
10,621