Building what customers will expect next. AI engineering, applied science & product leadership @Microsoft. Turning Frontier into trusted products at scale.

Cambridge, MA
It's the orchestrator that will continue to matter. Judgement in the form of experience from the driver of these models continues to bring value and importance, and that doesn't change anytime soon.
Unpopular opinion... Opus 4.6 (pre nerf) from a SWE was the goat. These models today while may be better at conceptual better complex architecture design, UI design, etc., however overly create complexity now, often take short cuts, don't fully implement or test, second guess and refactor, and respond back with difficult responses that take time to discern and figure out. With these models getting smarter and "better", I don't see SWE jobs getting any easier or faster or less complex.
2
106
📈Hill-Climbing is nothing short of rapid, impressive improvements, in very little time... It has been a true honor working with MAI and Office Science and Engineering to truly hill climb! microsoft.ai/news/hill-climb…
As @satyanadella says, we're making great progress on shipping MAI models that are higher quality, faster and cheaper across MSFT. In PowerPoint, our image model cut costs 84% compared with GPT-Image-2. In OneDrive, it lifted save rates 26% and cut latency by ~25%. It's also the default model in Bing delivering great quality and performance. In Dragon Copilot, our transcription model now covers 58 languages and halves the error rate on multilingual clinical transcription. And of course the flights we have in motion on GitHub and Excel will no doubt be huge too! More details in the blog here: microsoft.ai/news/introducin…
1
115
Louis Maresca retweeted
As @satyanadella says, we're making great progress on shipping MAI models that are higher quality, faster and cheaper across MSFT. In PowerPoint, our image model cut costs 84% compared with GPT-Image-2. In OneDrive, it lifted save rates 26% and cut latency by ~25%. It's also the default model in Bing delivering great quality and performance. In Dragon Copilot, our transcription model now covers 58 languages and halves the error rate on multilingual clinical transcription. And of course the flights we have in motion on GitHub and Excel will no doubt be huge too! More details in the blog here: microsoft.ai/news/introducin…
38
62
462
103,697
Louis Maresca retweeted
Apple accuses OpenAI of stealing trade secrets in a lawsuit that could reshape Silicon Valley alliances, threaten hardware ambitions and complicate OpenAI's path to an IPO, as Leo Laporte, @wesley83, @LouMM, & @NotPatrick discuss the fallout on This Week in Tech!
2
1
3
279
🇮🇹 After nearly a month in Italy and Sicily, enjoying incredibly fresh food, my body thanked me! No sugar jitters, bloating, swelling, or anything else. Now back to eating USA food and experiencing the same strange digestive issues. It’s the result of what our country allows in our food and what we permit agricultural companies to do to produce it (i.e., biologically and chemically)! Is it too late for the US, or can we turn the tide and fix this epidemic?
1
5
494
Want to use custom experiences in Copilot? Skills are where it is at!
Today we’re bringing skills to Copilot for Excel, giving teams a new way to scale their expertise across every workbook.
1
1
129
Louis Maresca retweeted
Such a privilege to work with Microsoft to bring claws to enterprises!
"You can run OpenClaw inside your company now." Annoucing our work with @Microsoft to bring OpenClaw to the Microsoft and Windows ecosystems. Claws now work securly in the enterprise.
132
309
4,693
479,287
Louis Maresca retweeted
Super excited to announce seven new world-class MAI models today. They represent what we consider a new era in AI designed to keep you in control and on the frontier. First is our text foundation model, MAI-Thinking-1, exceptionally strong on reasoning and SWE tasks. - It’s a 35B active parameter MoE with a 256K context window. Independent human raters on Surge prefer it for overall quality in blind side-by-sides versus Sonnet 4.6, and it’s achieved 97% on AIME 2025, the key measure of its general-purpose reasoning abilities. - It's at 53% on SWE Bench Pro, placing it right alongside Opus 4.6 on one of the toughest coding benchmarks. - And since we co-designed our models with our own silicon, MAI-Thinking-1 is optimized on our MAIA 200 chip. Benchmarking head-to-head against the GB200, we see 30% better performance per dollar as well as a 1.4x performance-per-watt gain when running our MAI models on the MAIA 200 end-to-end. Next is MAI-Image-2.5 and its Flash variant. Two super strong models now at #2 on the leaderboards, surpassing the score of Nano Banana 2 on image editing. Last for now is MAI-Code-1-Flash, our new inference efficient coding model, especially tuned for VS Code and GitHub Copilot CLI. - Code-1-Flash achieves 51% on SWE Bench Pro, despite having just 5B parameters, putting it closer to Haiku in size but cheaper in cost. All of this is the foundation for Microsoft Frontier Tuning. It lets you customize our models to create custom, company-specific agents that only you control. You can make our model, your model. Your data. Your agents. Your moat. Early adopters are already seeing a difference. When we tuned our models for McKinsey’s tasks, MAI delivered the highest win rate, outperforming GPT-5.5 on quality, while being 10x lower on cost. Also really excited to be collaborating with the amazing team at Mayo Clinic to jointly train a new frontier AI model for healthcare. Our announcements today mark another milestone on the road to humanist superintelligence. You can learn more and about our other new models in our latest blog: microsoft.ai/news/building-a…
193
523
3,778
1,323,543
Louis Maresca retweeted
Excel has quietly been Turing complete for a long time. Nice to see it now edging toward "AI complete"—SGD, attention, next-token prediction… all in cells.
Excel Copilot one-shotted a tiny GPT-style language model for me inside a spreadsheet: embeddings, causal attention, weights trained with stochastic gradient descent, next-token prediction, and a live slider to watch it learn.
126
130
1,180
232,962
Great Episode! @GlennF is back! Plus the insights and wit of @wesley83 is always refreshing! Love being on @TWiT and the big show, it always gets adrenaline going. 👏 twit.tv/shows/this-week-in-t…
Anthropic releases a new Claude model while conceding it trails an unreleased one, Snap lays off 16% of its staff, and Live Nation loses its monopoly case on This Week in Tech with Leo Laporte, @LouMM, @wesley83, & @GlennF!
2
73
Superman Day slips in like a quiet promise from the skies. Even when the world feels lost in shadows, the oldest heroes still remind us that light, hope, and goodness never actually quit. Superman and Lois Lane, ACTION COMICS #1 in 1938, thanks to Siegel and Shuster.
1
76
Thoughtful post. It isn't doomsday (homage to Superman Day 🦸), but an evolution of the world's economy in two years. How will your universe change because of it?
I told you that Anthropic believe that 50% of jobs will be done by AI in about two years, which was the average of what its own AI researchers believe. The entire staff was polled. Some believe AI leaders should keep that secret. I believe that we should be honest. So strongly disagree. But now there has to be real leadership on jobs. @DarioAmodei should join @elonmusk who is showing the best leadership. See my previous post for how.
5
1
112
Opus 4.7 - It's creeping up to level📈creepy. Another Thursday release... where's OpenAI?
Introducing Claude Opus 4.7, our most capable Opus model yet. It handles long-running tasks with more rigor, follows instructions more precisely, and verifies its own outputs before reporting back. You can hand off your hardest work with less supervision.
71
Here comes the announcements...before Friday...
It’s finally here! 🚀 Huge thanks to the @opencode team for the seamless integration. Qwen3.6-Plus and Qwen3.5-Plus are now live in Go! Update now to try it out! 👇
54
🥜@UberEats you're only as good as your weakest link your system and your customer service. “Add item to order” charges you all of the long range delivery fees, instead of just actually adding it to an order. A $4 item becomes a $18 item. immediately canceling, charges a $5 fee.
4
1
4
171
The icing on the cake is that you can’t get that money back; they won’t refund it. They refuse. They read from a script. customer service, zero. ordering system , zero. how many times I will order from Uber ,including rides? zero! good work Uber.
2
1
102
It keeps getting better @Uber @UberEats @dkhos It keeps getting better and better. “Quality,” they say. “Improving,” they say! Blame the customer is GREAT start! 👏
1
50
Lyft for me. Sold Uber stock, i have zero confidence in corporate operations
81
Transcribe is fabulous. Take your recordings and create highly accurate transcriptions. Honestly, those expensive transcription services have to watch out. plus, the Voice model is awesome!
Three models. Three top-tier results. All shipped within just a few months by the @MicrosoftAI team. - MAI-Transcribe-1 dropped today, the most accurate transcription model in the world across 25 languages according to FLEURS WER benchmark. - MAI-Voice-1 sets a new standard for natural speech. - MAI-Image-2 lands as a top 3 model family on @arena. We've been building with them - now you can too. All 3 available now on Microsoft Foundry.
1
112
If you haven't already, try it out! msi-playground.microsoft.com…
40