The open platform for automating development. Infrastructure to build, measure, and interact with agents across the SDLC

Pinned Tweet
Introducing Scorers: Agents that grade your agents. Use LLM-as-a-judge to grade past coding agent sessions on: • Quality • Efficiency • Compliance • Or any custom dimension Scores feed into performance measurements and automatic self-improvement for software factories
15
11
181
480,253
Warp retweeted
We @warpdotdev are teaming up with @warpdotco to host a poker tournament this Wednesday evening in NYC! Luma link & buy-in details in thread 😁
11
3
24
5,820
Grok 4.7 has landed in Warp and the Warp Agent CLI. Connect your @grok subscription to get started
3
2
23
7,167
Claude Opus 5.5 is now available in the Warp Terminal and the Warp Agent CLI.
1
2
48
4,839
GPT 6 Sol and Luna are now available in the Warp Terminal and the Warp Agent CLI
2
28
3,625
Introducing Scorers: Agents that grade your agents. Use LLM-as-a-judge to grade past coding agent sessions on: • Quality • Efficiency • Compliance • Or any custom dimension Scores feed into performance measurements and automatic self-improvement for software factories
15
11
181
480,253
Finally, use these metrics to have agents suggest improvements to your setup automatically. Self-improvement agents run on a schedule and review failing grades to suggest changes to your setup. Here is an example agent skills PR with evidence cited from previous scoring runs:
1
1
861
If you want to set this up, scoring is part of Warp Factories in early access. We are offering up to $10k in usage to qualified companies. warp.dev/factories/request-a… 🔖
1
826
In a fully humming factory, the key characteristic is that it’s a closed loop, measurable, improvable system. This should be the goal. In such a system everyone is working from the same context, in public, in a fully-audited, observed way. Agents themselves are observing the skills and config that drive the system and suggesting improvements. Platform engineers are able to extend the system to integrate it into all internal systems. Engineering leaders can see productivity metrics and understand what changes are being made to improve them. The whole thing is running empirically, not on vibes.
7
7
49
189,686
It may be time to start automating some of those Claude Code tabs you have running...
The software factory approach is getting pretty popular, but it can be daunting to adopt all at once. Let's go through “crawl, walk, run” steps for making the transition from local, interactive agents to automated cloud development.
Article

Adopting the software factory model: crawl, walk, run

The software factory approach (a closed agentic loop that runs in the cloud) is growing in popularity, but it can be daunting to adopt. In this post I’ll go through “crawl, walk, run” steps for making

5
12
3,897
Warp retweeted
Skill GitHub has been achieved internally
1
1
30
8,584
Warp retweeted
/remote-control with Grok Build CLI is my favorite feature of this launch. Run `grok` in the Warp Terminal and click `/remote-control` to copy a session sharing link to your keyboard. With this, you can: > paste it in your browser and continue the same exact session from there > send the link to a friend to work on (v helpful for hackathons, etc), your phone, or another device > embed the link in an iFrame anywhere (lots of cool project i'd love to see doing this) and did i mention... this + all of Warp is open source? 🙂
Warp now has built-in support for the Grok Build CLI. - Use Warp's rich input for agent prompts, with support for longer pasted prompts and multi-cursor - Use /remote-control to share your agent session to another device - Access the file explorer and code review panels
2
4
37
9,450
Warp retweeted
Model selection is a huge driver for AI spend; we just reduced our cost-per-PR from $80 to $30 by switching from Opus 5 to GPT 5.6. We made this decision not by guessing, but by benchmarking. I want to show you how you can do this too! Join me for a live session on building custom model benchmarks: Thursday, September 17th 2pm ET We'll build a custom benchmark that replays your past agent conversations against a set of different models to find the best cost/performance pick. I'll whiteboard the entire setup so you can do this yourself, and we'll walk through the batteries-included version with Warp Factories. RSVP: luma.com/warp-8vtl 📅
4
2
16
4,382
Warp now has built-in support for the Grok Build CLI. - Use Warp's rich input for agent prompts, with support for longer pasted prompts and multi-cursor - Use /remote-control to share your agent session to another device - Access the file explorer and code review panels
27
5
162
930,057
Software factories should be built on an infrastructure stack that is open, composable, and defined-in-code. This article lays out the principles that apply to anyone who is looking to move to a factory model.
Article

The software factory stack

Software factories should be built on an infrastructure stack that is open, composable and defined in code. Open == works with any model, agent and hosting configuration Composable == you can adopt

31
14
337
33,310
Our cost-per-PR keeps trending down. Now we're at $30. The biggest drivers: - Benchmarking models on our own data to find the most efficient cost/performance pick. GPT 5.6 Sol (high) today - Letting agents review past conversations and PR skill patches for bad instructions and adjusting our skills accordingly $30 still seems expensive on average though. I think we can get this to $10 or lower, and we are still figuring out what good looks like here. Obviously not all PRs are equal.
15
4
70
6,989
Our company’s cost-per-PR dropped from $80 to $30 by switching to GPT 5.6 Sol. We built a benchmark to replay our team's agent runs across model providers. GPT 5.6 Sol got the highest code quality at a 66% lower cost vs. our old default (Claude Opus 5).
9
4
61
7,349