Moritz Stephan retweeted
We are looking for a founding designer to help shape the product at @hone. If you are this person or you know anyone we should talk to send me a DM!
52
3
219
12,195
we're proudly optimizing business metrics @chrisbarber appreciate the shoutout!
I made a list of great startups to join. It's called the Breakout List. The list has 92 companies. These are the 20 with 25 or fewer employees: - Hone (@moritz_stephan, @CarloWillem, @oqbrady) - Normal (@ansonyuu, @hudzah) - Standard Intelligence (@G413N, @devanshpandey) - Tacit Labs (@ninklefitz, @AmDroste) - American Terawatt (@atroyn, @rslparker, @aranibatta) - Conduit (@clemvonstengel, @riopopper) - Convergent (Omkar Savant, Vivek Katara, @debnilsur) - Core Automation (@MillionInt, @_arohan_) - Engram (@dan_biderman, @EyubogluSabri, @realJessyLin) - Instinct (@noahrshinn) - Keenable (@styskin, Matthias Petri) - Lumaril (Mark Elliot, Ben Duffield) - Neion Bio (@Dimkell, Sam Levin) - Pangram Labs (@max_spero_, @bradley_emi) - Quadrillion (@echinaceous) - Re (@karnsaroya, @AnandDhillon, @thecliffwhite, @benaneesh) - Ricursive (@annadgoldie, @Azaliamirh) - Sail Research (@neilmovva, @blintzbase) - Trajectory (@rronak_, @michaelelabd, @QuantumArjun) - Watney Robotics (Sean Cheong, Ryan Gannon) Picks from Elad Gil, Charlie Songhurst, Keith Rabois, Mike Vernal, Alana Goyal, Sonya Huang, Ramtin Naimi, Marc Bhargava, Cory Levy, Aashay Sanghvi, Konstantine Buhler, John Luttig, Varun Gupta, Ray Tonsing and Avichal Garg. Disclosure: I'm a small investor in American Terawatt, Convergent, Standard Intelligence and Trajectory (in this post), and in Factory, Physical Intelligence and SF Compute (elsewhere on the list). I didn't vote. The full list is on Breakout List.
2
3
66
9,316
Huge progress by @silasalberti and team 🔥
Introducing SWE-2, our closest model yet to the frontier. On leading evals, it scores on par with recent frontier models – at up to 70% lower cost. We scaled RL to multiple trillions of parameters, with a refined recipe that pushes the Pareto curve on both capabilities & cost.
3
1
81
4,997
If OpenAI can throw ~$10M worth of frontier tokens at solving Navier Stokes, what has to be true so enterprises can do the same to solve their most critical business outcomes?
12
1
95
9,528
welcome to the team @whosmatu 🚀
Happy to announce that I've joined Hone. We're a small, ambitious team taking a new angle on what AI working for you will actually look like. Models have advanced incredibly fast, but 90% of the tools the world uses at work haven't kept up. We're changing that.
2
1
58
6,124
Moritz Stephan retweeted
👀
1
15
3,914
Moritz Stephan retweeted
Today we're announcing our $40M Series A at a $400M valuation, led by @a16z , with participation from existing investors @8vc, @pearvc, and @BloombergBeta and new investors @HRTVentures and @nextladder. Alongside the fundraise, three more announcements: - Vals Smith: is now generally available. Anyone can create a custom coding benchmark from any GitHub repo with 120 free credits to get started. - Frontier Risk Benchmarks: We are releasing the RSI Index in collaboration with @CoreWeave and just launched ReverseEngBench, a new cyber benchmark built with Columbia University, Tufts University, UC Berkeley, and UCLA. We are also sharing our initial work in mental health, with more to come across environmental impact, military, and biosecurity. - Website + Vals Index 2.0: We completely rebuilt the Vals website and have released Vals Index 2.0, with coverage of more of the economy. Our revenue has already grown 8x compared to all of 2025. Our customer base doubled and the team tripled in 6 months. Our results have been cited in model cards from OpenAI, Anthropic, Google, Meta, and xAI. The AI economy runs on self-reported grades. When a model ships, the scores come from the company that built it. No other trillion-dollar industry works this way. Finance has ratings agencies. Medicine has the FDA. AI has vibes and vendor benchmarks. We built Vals to be the independent evaluation layer the industry is missing. Check out our new website and try Vals Smith!
76
42
366
242,634
RT @_mnat_: When life gives you codex and free will Of course you make yourself a massive LED Jumbotron in your office to monitor everyone…
1
1,443
We are launching Hone today. Over the last few years I had a front row seat to how AI has reshaped software engineering. Model progress exceeded my wildest expectations. Yet outside of engineering, most organizations struggle to derive value from AI. At Hone, we enable organizations to direct AI at their most complex business problems: AI that owns outcomes over weeks and months, not tasks. Read more about our mission at hone.com. I am immensely grateful to @ScottWu46 , @stevenkplus1 , @walden_yan , @russelljkaplan and all the friends at @cognition for the incredible journey and everything you taught me. Thank you for your support in this new journey.
Announcing Hone Intelligence has become abundant. Yet the world looks remarkably similar to how it did five years ago. With every model release, the gap between what frontier AI can do and the economic value derived from it widens. Closing the gap requires re-organizing work around organizational outcomes, not individual tasks. Hone builds AI that creates, orchestrates, and improves agents and software continuously to own organizational outcomes over weeks and months. We are ex-founders and early core contributors to Cognition, Mercor, Ramp, and OpenAI. We obsess over real-world value, not theoretical benchmarks. Our core beliefs on closing the gap in the thread below.
70
26
410
73,689
Moritz Stephan retweeted
Introducing Vals-Smith: turn your code base into a customized benchmark. Public benchmarks tell you which model is strongest overall, not which model is the best on your code. Vals-Smith turns your merged pull requests into real coding tasks and measures the percentage a model can actually resolve. New models ship every week. Vals-Smith tells you which one to trust with your code.
30
33
296
46,614
Moritz Stephan retweeted
Genius AI is bringing intelligence to the physical economy, where some of the most valuable and human work in our daily lives takes place. Since Day 1, we have been obsessed with helping talented practitioners do what only they can do as technology takes on the repetitive tasks. We’re excited to share that we raised a Series D at a $1.15B valuation to accelerate our vision. We are also announcing a new brand, Genius AI, as our platform and customer footprint have evolved. Alongside our core flagship GlossGenius product, we’re releasing products that run more of the admin work on their own, so business owners can finally do what they dreamed of. To the 125,000 businesses who have trusted us to save them 112M hours of admin work and earn billions of dollars in extra revenue already, thank you for your partnership. Our work is just beginning as we make your most repetitive workflows disappear. Grateful for our investors, including @Lux_Capital that led the round with participation from @BessemerVP @ImaginaryFund @2048vc @L_Catterton and @StepStoneVC . Here’s a note from our founders @DCohenShohet @Leahcs with more detail below ⬇️
46
26
132
110,959
Moritz Stephan retweeted
Introducing Aidan, your newest AI employee Aidan is the first computer-using AI built for realtime conversation We’ve raised $45M from @sequoia and @8VC to bring Aidan to the world @NotionHQ and @DecagonAI already use Aidan to run customer interactions. This is how @withsableai works:
300
143
1,541
737,560
Moritz Stephan retweeted
1/ We’ve raised over $1B at a $26B valuation, led by @Lux_Capital, @generalcatalyst, and @8vc. Our enterprise usage has grown >10x since the start of this year, and our run-rate revenue grew to $492 M. We launched Devin two years ago as the first AI software engineer. Since then, cloud agents have gone from niche to mainstream, and today they are the fastest growing way to create software.
164
193
2,458
919,745
agents should feel like a co-worker, not a tool. super excited about this launch
Introducing Devin Auto-Triage: Your AI first-responder with long-term memory. Devin can monitor incoming bugs, alerts, and incidents, investigate them, and come back with context, next steps, or a PR.
5
4
102
26,806
Moritz Stephan retweeted
I am in love with Devin from @cognition , now i know why it was priced at $500.
1
4
49
6,452
it was a blast working with @spdling, @rhythmrg, @raymondmfeng and the rest of the @appliedcompute team. Splitting capability maximization (i.e. be good at finding bugs) and product alignment (i.e. short rollouts) into two distinct phases while training made a big difference here and can be useful for other specialized models when real-world product constraints matter
Today we're releasing SWE-check, a specialized bug detection model we RL-trained with @appliedcompute that matches frontier performance on internal in-distribution evals and makes meaningful progress on out-of-distribution evals, all while running 10x faster.
2
45
7,741
devins manage devins manage devins manage devins manage devins manage devins manage devins manage devins manage devins manage devins manage devins manage devins manage devins manage devins
Devin can now manage a team of Devins. Devin will break down large tasks and delegate them to parallel Devins that each run in their own VM. Over time, Devin gets better at breaking down and managing tasks for your codebase. Available now for all users.
2
2
40
6,680