We build models for agentic coding and long-horizon tasks. Try Laguna: poolside.ai/get-started

San Francisco, CA
Pinned Tweet
Today, we’re releasing Poolside Desktop Assistant. One place to run coding agents across macOS, VS Code, and Visual Studio. We built it for ourselves and have used it every day for the past year. Now we’re opening it up to everyone.
21
51
326
80,098
Poolside retweeted
Frontier-Bench looks good. However at n=74, it may be too small to reliably distinguish model checkpoints. Our quick analysis suggests 100+ tasks would provide a more reliable signal...
We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work. Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort. Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%
2
2
8
1,369
We’ll be at ModCon ’26 in SF on August 18! @varunrandery is joining the frontier models panel to talk about building models for long-horizon agentic work, where open models differentiate, and what it takes to get them into production. Join us in SF or tune in online ↓
Building frontier models involves a hundred decisions rarely discussed publicly: when to stop scaling, what data to keep, how to differentiate, whether to release the weights. On August 18th at ModCon '26, Paige Bailey of @GoogleDeepMind, Joseph Spisak of @reflection_ai, Victor Su-Ortiz of @MiniMax_AI, and Varun Randery of @poolsideai are talking through the challenges of the industry from four different vantage points. Register for the livestream: luma.com/modcon-livestream
1
2
23
5,651
Agentic evals are messy. A benchmark score tells you something about model performance, but it also reflects the whole system around it: the harness, sandbox, dependencies, timeouts and sometimes a loophole the agent found in the task. That’s why trajectories matter so much to us. They show what the agent actually did and whether the score means what we think it does. We publish them to make that evidence transparent and auditable, giving the wider community more to learn from. Watch @aalSonOfRavi and @ConnorBAdams go deep on all of this with @petergostev from @arena, including some surprisingly creative reward hacks!
Meet… Arena Conversations. @petergostev sits down with @poolsideai researchers,Connor Adams and Aalhad Patankar, to discuss how they’re building frontier open coding models. Episode drops today at 10am PT on our YouTube. They dig into Poolside's Laguna model family, why the team publishes full trajectories (not just benchmark scores) for anyone to audit, and how "experiments" at Poolside range from data mixes to harness design to reward-tuning decisions run tens of thousands of times a day.
2
8
67
8,789
Poolside retweeted
another 2 weeks of Laguna S 2.1 Free on Nous Portal, thank you @poolsideai
The new Laguna S 2.1 model by @poolsideai is now free for 2 weeks on Nous Portal. At 118B total parameters with 8B active, it's quick to run and the most capable model they've released so far. Try Portal today at portal.nousresearch.com/sign…
11
13
178
27,310
Come hear @mgalle and @sudip_r0y get into the weeds on post-training, agent RL, continuous learning, and what it takes to make models work reliably in production!
First episode of Field Notes, our new monthly series on the leaders shaping AI. @sudip_r0y, Adaption Co-founder, and @mgalle, post-training lead at @poolsideai discuss the last 5% of reliability and what it actually costs.
2
22
5,979
we knew the community could make Laguna faster. 2.6x faster is pretty fun :) huge congrats to the winner, and to all 35 solvers who spent the last few weeks pushing Laguna XS 2.1 further! more reasons to build in the open, together 🤝
We started the MLX.fast challenge with @poolsideai with one key question. How much faster can the community make Laguna XS 2.1 run on a Mac? The answer so far has been 2.6x faster. Nearly 1,800 submissions and 35 solvers later, the challenge comes to a close today. And there’s one more thing. One of the solvers is taking home a Mac Mini M4 24GB! Winner below ↓
8
15
163
13,049
Poolside retweeted
Our SENPAI agent is currently no.1 on the @poolsideai x @eigenlabs inference optimization comp (for now) and first to break 200 TPS decode Lovely first validation after a complete re-write of the agent to use @OpenHandsDev instead of cc
Excited to launch MLX.fast with @eigenlabs today. It's an open autoresearch competition to make Laguna XS 2.1 inference as fast as humanly (and agentically) possible on consumer Macs. Eigen's agents already found 36.8% faster inference, and that's before the competition even started. The best part of open weights is that the community takes a model further than any of us could. 
Can't wait to see what everyone does on the leaderboard!  mlx.fast/
5
6
19
4,409
Poolside retweeted
Poolside are pioneering models built specifically for local hardware. Laguna S 2.1 is a great model for DGX Spark / MacBook. The number of tokens generated is probably an order of magnitude more if you include tokens generated locally.
We just crossed 10T tokens served across all our models in less than 3 months! Laguna S 2.1 is pushing new highs at ~300B tokens a day, and has processed +2T in the 14 days since release. That’s across @OpenRouter @vercel and our direct API. Really good to see demand for open models keep accelerating.
2
25
4,269
We just crossed 10T tokens served across all our models in less than 3 months! Laguna S 2.1 is pushing new highs at ~300B tokens a day, and has processed +2T in the 14 days since release. That’s across @OpenRouter @vercel and our direct API. Really good to see demand for open models keep accelerating.
14
19
183
17,145
Watch Laguna S 2.1 climb at poolside.ai/pulse *Pulse currently shows OpenRouter traffic only, the rest is distributed across Vercel AI Gateway and direct API.
4
1,369
Poolside Desktop Assistant has been out for a week, and we’ve received so much great feedback! Today we’re releasing version 1.4.0, with a bunch of new features and fixes based on your bug reports: - Native steering and queueing where supported - Proper plan mode and agent Q&A - First-class subagents, with full transcripts for Claude and better status reporting for Codex - Much faster inference for local models - A whole load of smaller bug fixes More below!
4
9
81
9,856
We made local inference faster for any model you run in Desktop. Try it with Laguna XS 2.1. Tabs and split panels are more responsive too, and code files open almost instantly.
1
4
712
Thank you to everyone who tried Desktop and shared feedback in our first week! A lot of 1.4.0 came directly from you. Update the app, mix and match any model with any harness, and have fun with it. Get started: poolside.ai/get-started Then come tell us what we should build next: discord.gg/NAnCKRbZa
1
5
789
Poolside retweeted
One of our engineers asked @poolsideai's Laguna S 2.1 to transform a 715-file C++ game from neon cyberpunk into an Ancient Greek aesthetic. It orchestrated 3 different models, refactored the code, and produced a playable build. baseten.co/blog/laguna-s-21-…
4
2
33
3,557
Poolside retweeted
Made some improvements to bcode + local models Biggest winner is Laguna S2.1 which was compromised by provider issues, it gained +23% score since last post, now is very competitive
3
17
2,231
Poolside retweeted
excited to share we have made Laguna 2x faster on consumer Mac machines!! i want to thank all the participants of this challenge -- this wouldn't have been possible without you. all this with no speculative decoding; we're going to introduce it soon -- we want to make sure we ship an anti-hack verifier for our system!
9
10
83
10,113
We've improved our serving efficiency and increased rate limits by +10x. The update is live on @OpenRouter @vercel AI Gateway and platform.poolside.ai. Thank you to everyone who has used the model, shared feedback and stuck with us while we improved the experience! Usage is already climbing. On OpenRouter alone, Laguna S 2.1 is on pace for ~250B tokens today, 4x our daily average this week. We are also taking 10% off our paid endpoint on OpenRouter. It's a dedicated deployment with the full 1M context window for the best performance on harder tasks. Run Laguna S 2.1 in pool or Poolside Desktop Assistant, or plug it into @opencode, @NousResearch Hermes Agent, @kilocode, @cline or @pidotdev and let it run over the weekend. We'll be watching the graphs.
11
11
148
25,187
We also identified and fixed the looping-in-thinking issue some of you experienced with Laguna S 2.1. This was due to an interaction between DFlash and TensorRT-LLM which we are investigating, as well as default serving at an incorrect temperature. Our recommended temperature is 1.0. Please keep the feedback coming. We’ll share a deeper technical write-up on what we found soon.
8
8
114
26,797
Special thanks to @antirez @CardilloSamuel @dealignai @sudoingX @JoelDeTeves @ivanfioravanti @Blackwellboy @0xSero @onusoz and many others who have been relentless in testing Laguna S 2.1 and sharing feedback!
7
3
52
12,874