President, New Enterprise at Cognition, Prev. Windsurf CEO, AI @ jeffwang.substack.com

San Francisco, CA
Pinned Tweet
We’ve raised $2B at a $48B valuation, since our last round a few months ago, we've gone from $492M to near $900M, the era of self-driving software has begun Devin has gotten even more awesome in this same time period. From adding security swarm, Devin Automations, auto-triage, iOS/Windows/Android support, and a whole assortment of other features, the experience feels more robust than ever. Our customers are our best product managers, we work very closely and deeply understand their business and problems, and implement new features and fixes faster than ever It’s been a privilege working with some of the most talented people on the planet. I'm humbled to be on this ride and look forward to continue building the future of software engineering
The world needs far more software than it can build. Cognition exists to change that. We’ve just raised over $2B at a $48B valuation, led by a16z, Accel, Founders Fund, General Catalyst, and Avenir. Since our round in May, run-rate revenue has grown from $492 M to almost $900 M.
13
3
182
11,778
We’ve crossed $1B in annual run rate This is an amazing milestone but we owe it to our customers. We work together with them and our team on the most challenging problems and it shows in the product and how we collaborate Honored to be a part of this moment, but there’s much more work to be done
Cognition has crossed $1B in annualized revenue run rate. This milestone belongs to our customers. Here's how a few of them are building with Devin.
5
1
90
3,734
So today we didn’t have a new Devin feature launch so instead we launched a new website
Devin has a new home page.
3
2
55
2,601
Devin now natively works with Microsoft Teams and Microsoft 365 Now you can DM Devin, set automations to monitor channels, or give access to your MS apps like email and calendar Automations are a personal favorite, now Teams users have access to it!
Devin loves building apps on its own PC. Now, it works even better on yours. Introducing two updates to Devin for Microsoft users: native support for Teams, and a first-party Microsoft 365 integration to connect to your mail, calendar files, chats, and more.
1
26
1,720
Today the Pareto curve got updated with new Anthropic Opus 5.5 and OpenAI GPT-6 Sol and Luna models. On the cheaper end GPT-6 Luna is the best for its low end price and Opus 5.5 is the best bang for buck with frontier performance. We are still seeing signs of Fable 5.1 being preferred internally and we are aware benchmarks aren’t perfect, but this sets the new bar for models on the curve It’s also worth noting that other types of tasks are still model dependent as well, we prefer GPT-6 Astra for computer use, as an example, and SWE-2 for model pairings for Devin Fusion since it has been trained to be cost effective across the curve
9
6
78
6,322
Jeff Wang retweeted
GPT-6 Sol and Luna are now available in Devin. On FrontierCode 1.1, GPT-6 Sol matches GPT-5.6 Sol’s score at 61% lower cost per task. GPT-6 Luna scores above GPT-5.6 Luna at about a quarter of the cost. At under $0.10 per task, it is the cheapest model on the leaderboard.
55
57
751
384,982
Welcome @bcamhi to Cognition! He’s already been cooking
I’ve joined @cognition to lead marketing! Looking forward to sharing what we’re cooking up. We’re hiring across the board and we’re going to build a legendary marketing team…my DMs are open :)
1
32
1,571
For those that still prefer using CLI but miss the benefits of using cloud, we’ve added /cloud to change to a cloud session and /handoff to bring it back to local
Introducing Devin Cloud in Terminal and devin ssh: two new ways to use Devin's computer. Create, steer, and resume Devin Cloud sessions right in your CLI with /cloud. And for the first time, SSH into Devin's dedicated VM, then /handoff the work back to your own device.
3
1
64
3,529
Lots of companies are rolling out model routing but many underestimate how good Devin Fusion is. It turns out handing off to a dumber agent doesn’t work well, because an agent doesn’t know if its work is subpar or good When using a combination of Astra+Fable and SWE-2, having frontier models monitor cheaper models for certain tasks need to run in parallel (main+sidekick) for this very reason We pick the smarter model which can then keep the cheaper models quality high, this also further saves costs because you can increase cache hits For this very reason Devin Fusion has become a daily driver and I suspect it will for others (who want to optimize on cost without tradeoffs on performance)
15
9
104
5,760
Introducing Code Scans, our second product from agent swarms. You can scan across your code base for tasks such as making things more optimized, auditing outdated libraries, really anything you need thorough scanning for. Then with agentic map reduce you can then summarize the findings and changes to keep everything neatly organized
Some engineering tasks require reasoning across the codebase: Which code is safe to delete? What queries are slowing performance? Introducing Code Scans: codebase-wide audits for any goal. Devin investigates, reports findings, and opens the PRs. Powered by Agentic MapReduce.
8
4
76
4,673
This is actually a big deal, you can now spin up Mac VMs with Devin. That means if you want to build and test iOS apps you can do so completely end to end from your slack or web interface. One of the other big customer use cases is if someone reports a mobile app issue Devin can automatically reproduce, fix, and send you proof of the issue, demo of it working and of course a working PR
Special delivery: Devin just got a Mac 🍎 Now Devin can: 1. Build & test apps on its own Mac VM with iOS simulator 2. Send a screen recording via Slack 3. Send a TestFlight link so you can start using it 📲
13
13
210
31,881
While it wasn’t initially obvious yesterday with our SWE-2 launch, it turns out when you optimize a model for cost and performance, it makes for a great model to pair as a sidekick to Devin Fusion Check out our evals with @ArtificialAnlys and @ValsAI to prove that you can get near frontier performance for a 36-39% discount You can now use Devin Fusion in CLI, try this out on your own agents!
Introducing Fusion in Devin CLI The most efficient frontier harness for Fable & Astra; 39% cheaper across coding benchmarks. Pick your favorite model for planning and a cost-effective model for execution.
6
55
3,704
Oh yeah we also released Devin Voice today
Introducing Devin Voice 🦦 ☎️ Your favorite AI software engineer just got a landline. You say it, Devin ships it. Powered by GPT-Live and our new SWE-2 model.
1
4
66
3,511
Jeff Wang retweeted
Welcome @jkelleyrtp and the entire Dioxus team to Cognition! Dioxus is one of the most beloved open source Rust frameworks. We're proud to continue supporting Dioxus, Blitz, Taffy, and Subsecond while bringing the team's expertise to Devin's VM, computer use, and testing.
We have some very big news to share. 🎉 Today, Dioxus Labs is joining @Cognition to accelerate the development of Devin, Cognition’s autonomous cloud coding agent. We are excited to continue building Dioxus while also helping Cognition shape the future of software engineering.
20
24
294
44,202
Today we're releasing SWE-2, our latest model Not only does SWE-2 come within 1 point of Fable 5.1, it is also 64% cheaper as well. While it's awesome that we're on the pareto frontier, the more impactful outcome is what our research team has innovated on: -Using cost penalty as a reward function, it's even cheaper than SWE-1.7 -RL optimizations, we reduced memory usage despite the model being 3x the size of SWE-1.7 -Tripling the number of RL environments and scaling up All these advances have enabled SWE-2 to drastically move Kimi-K3 to the Pareto Frontier. We've also applied our trustworthiness alignment to the model to reduce bias and security risks. I'm very proud of our research team for this massive accomplishment! SWE-2 is now live in Devin Cloud and Devin Local agents, we have some great promos going on right now for using SWE-2 so please let me know what you think
Introducing SWE-2, our closest model yet to the frontier. On leading evals, it scores on par with recent frontier models – at up to 70% lower cost. We scaled RL to multiple trillions of parameters, with a refined recipe that pushes the Pareto curve on both capabilities & cost.
28
8
256
12,043
Jeff Wang retweeted
here's the writeup of the factorization! cognition.com/blog/factoring…
Last week we published a factorization of RSA-260. Today, we’re sharing the methodology of how Devin and a Cognition researcher built the world’s fastest GPU optimized lattice siever, to make factoring numbers 10x cheaper than the previous state of the art: cognition.com/blog/factoring…
15
65
694
233,225
Jeff Wang retweeted
GPT-6 Astra is now available in Devin Desktop and Devin CLI!
GPT-6 Astra is coming to Devin. On FrontierCode 1.1, Astra performs within 0.4 points of Fable 5 at a 64% lower cost. It also sets a new SOTA on our internal testing benchmark, generating more comprehensive tests, clearer reports, and better video evidence.
21
23
187
73,262
GPT-6 Astra is looking really good for cost per task, identical to Fable but 64% cheaper We are about to see Devin Fusion get a lot cheaper as well
GPT-6 Astra is coming to Devin. On FrontierCode 1.1, Astra performs within 0.4 points of Fable 5 at a 64% lower cost. It also sets a new SOTA on our internal testing benchmark, generating more comprehensive tests, clearer reports, and better video evidence.
1
2
23
2,251
Getting on the cost/performance pareto curve is getting extremely competitive right now. If you are not on it, you are in danger of your model being unused UNLESS of course, the model can specialize in specific types of tasks. This is where we can plug Devin in for our Fusion and router swapping Having access to all models and understanding what they are good at is key here, in addition to our overall evals:
4
1
48
2,914
Thanks for @cl571128 for keeping frontiercode up to date
2
306