Command your own software factory.

San Francisco
Building on the release of MacroscopeBench, we’ve been researching post-training our own models at Macroscope, leveraging training infrastructure from @FireworksAI_HQ . Our co-founder, @Rob_Bishop, and Vivek Chauhan, product lead for training at Fireworks, will be sharing learnings from the collaboration on Sep 29, 11am PT.
2
2
7
28,906
GPT-6 Sol scored 80.3, #2 overall, behind only Claude Opus 5.5 max. It is ahead of GPT-6 Astra at less than half Astra's cost. Against GPT-5.6 Sol it matches the old model's best score at half the price
1
1
134
GPT-6 Luna, while ranking at #13 using max effort, still delivers great performance for extremely low prices. It reached 93% of GPT-6 Astra's score for about 1/60th of the cost! It holds a very respectable f1 score due to its very high precision (89.1%), second only to Astra
88
There are some new top models leading our code review benchmark. 🥇 Claude Opus 5.5 — 82.3 (f1 score) 🥈 GPT-6 Sol — 80.3 🥉 GPT-6 Astra — 78.0 See full results here: macroscope.com/benchmark
4
6
15
3,737
Compared to Opus 5, Opus 5.5 is better, cheaper and faster at every effort level, usually around 1/5th of the cost and up to 4x faster
1
1
133
Claude Opus 5.5 at max effort takes #1 with impressive recall (80.6%) and precision (84%). The downside: cost and latency, being the most expensive model we’ve tested xHigh effort level still delivers very strong performance, very similar to GPT 6 Sol max at half the latency.
1
3
177
Production is the worst place to find out if your code reviewer is good enough MacroscopeBench measures it first, against real bugs that shipped, just not in your repo. It is now the only trusted benchmark for code review on @FireworksAI_HQ SII Index fireworks.ai/specialized-int…
1
7
14
64,241
we rewind to the commit that introduced each bug and ask whether a reviewer would have caught it there. it's never told a bug exists. see which models currently perform the best at macroscope.com/benchmark
1
4
133
Learn more about how we built MacroscopeBench and how it works here: macroscope.com/blog/macrosco…
2
6
424
We're hosting our first webinar on 9/22! Ivan from Macroscope team is walking through the four things Macroscope code review actually does to a pull request: review, gate, approve, & report. We'll show you how to use AI code review tools to get a PR from open to merged without a human in the loop. Register below 👇 luma.com/1fw8z3lt
1
10
603
Will you run your Murmur agents on an iPhone Duo?
1
2
10
903
Macroscope retweeted
“A single engineer might be capable of directing dozens of productive agents, yet in practice can only keep a handful moving at once.” Exciting early preview from @kayvz and the @Macroscope team, making it easier for engineers to direct fleets of coding agents in the cloud and for work to stay in motion.
1
2
14
3,446
Macroscope retweeted
murmur ASMR
2
20
3,823
Macroscope retweeted
As a “Member of GTM Staff,” I also ship a ton of code to production with Murmur internally. this is what my workflow looks like with cloud agents 🧵
2
2
5
504
Macroscope retweeted
Murmur has significantly improved my productivity and changed the way I work. Being able to trust agents to self-drive from implementation through verification and fixes has been a game changer. Super excited for this! Happy spawning 🚀
2
6
653