CTO of Turing, ex-CVP @ Microsoft

Redmond, WA
CEO Bench is a stop towards answering the biggest question in AI: can agents create real value inside real enterprises? Answering it takes benchmarks that capture what domain experts actually know. Top frontier model scores 41.3/100. Enterprise work is far from solved.
2
6
163
Excited to share that I've joined @turingcom as CTO. After two decades in AI research, I can't wait to push the frontier here; working with the top AI labs and partnering with enterprises to put it to work. Onward. 🚀
Turing welcomes Ece Kamar (@ecekamar) as Chief Technology Officer. Ece joins to lead Turing's technology and research strategy, helping steer the company through our next phase of growth as we scale across both frontier AI labs and Fortune 500 enterprises. We’re advancing the AI frontier by providing the human expert data, RL environments, evals, and benchmarks that define what agents can do, and also translating that research into reliable, adaptable agentic systems enterprises can depend on, with sovereignty built in from the start. We’re doubling down on software engineering, enterprise knowledge work, and frontier STEM. Ece comes from Microsoft, where she led the AI Frontiers Lab as Corporate Vice President, driving research from foundation models to agentic AI and building the tools, benchmarks, and open-source projects used across the field, including the Phi models and AutoGen. She co-authored the "Sparks of AGI" research paper and holds a PhD from Harvard. Welcome to the team, Ece. We’re excited for what’s ahead!
36
11
333
22,950
🚀 Fara 1.5 is now on arXiv! 🚀 The report is a must read for anyone that wants to learn what it takes to build a SOTA computer use model. Fara 1.5 establishes a new state-of-the-art for its size, outperforming larger cloud models like OpenAI Operator and Gemini 2.5.
Fara1.5 is here! The tech report just landed on arXiv. New SOTA for computer use agents of its size, and it competes with much larger frontier models. Paper: arxiv.org/abs/2606.20785
1
9
1,837
New work from our lab on Next-Latent Prediction and stronger world representations for agents, paving the way for more capable long-horizon reasoning and decision-making. Well done @jayden_teoh_ @manan_tomar @KwangjunA @edward_s_hu @Tea_Pearce @pratyusha_PS Akshay Krishnamurthy @riashatislam @lamblabtsinghua @JohnCLangford
Next-token prediction is myopic. What if transformers learn to predict their own next latent state? 🌠 We present 𝗡𝗲𝘅𝘁-𝗟𝗮𝘁𝗲𝗻𝘁 𝗣𝗿𝗲𝗱𝗶𝗰𝘁𝗶𝗼𝗻 (𝗡𝗲𝘅𝘁𝗟𝗮𝘁): a self-supervised learning method that teaches transformers to form compact world models for reasoning and planning. It also unlocks up to 3.3x faster inference via self-speculative decoding! 🚀
1
7
24
3,330
Ece Kamar retweeted
Sometimes less is more. In a world that's constantly changing, the challenge for always-on agents isn't just acting, it's watching and waiting. We built SentinelBench, an open-source benchmark for evaluating agents on long-running monitoring tasks in dynamically changing environments. Our results show that how agents wait can matter as much as the model itself, cutting costs by up to ~10x, often with event better outcomes. More below 👇
How patient is your agent? We’re releasing SentinelBench: web monitoring tasks across 10 synthetic apps, designed to test whether agents can watch, wait, and act when the world changes. Turns out “how you wait” matters. A lot. SentinelBench, a Benchmark for Long-Running Monitoring Agents - Microsoft Research
2
5
459
MagenticLite is officially live on GitHub, with models available on Microsoft Foundry. We’ve optimized the entire stack end-to-end to deliver a faster, more efficient agentic experience powered entirely by small language models.
1
3
890
What powers this end-to-end harness: Fara1.5: Our computer-use model family (4B, 9B, 27B). The 9B flagship nearly doubles Fara-7B on web navigation, setting a new SOTA for small computer-use models. MagenticBrain: A 14B orchestrator model that plans, codes, and delegates.
2
225
Most AI agent benchmarks measure task completion. Not whether the agent actually represented you. SocialReasoning-Bench fills that gap — testing agents in multi-party scenarios like scheduling and negotiation. Our key finding: frontier models do complete the task, but routinely accept bad deals instead of advocating for the user. To learn more: microsoft.com/en-us/research…
2
9
12
1,135
Are frontier models truly ready to act as our delegates? The AI Frontiers Lab is releasing SocialReasoning-Bench to measure if AI agents actually act in a user’s best interest. Results show even frontier models struggle with social reasoning and due diligence.
Using SocialReasoning Bench, we observed a stable pattern across models—agents execute competently, but fail to consistently improve the user’s position, even with explicit instructions to optimize for user interest. msft.it/6011vPOLF
1
1
11
1,177
So much amazing work to be done at the frontier of AI — come build it with us! 🚀 We're hiring Senior & Principal Research Scientists at @ms_aifrontiers. If you work on agentic AI, multi-agent reasoning, continual learning, or synthetic data — we want to hear from you.
🚀 Hiring: Research Scientists 🚀 We're hiring Senior and Principal Researchers in Agentic AI at @ms_aifrontiers Lab at @ResearchU Our focus is on developing self-improving agentic systems, agents that learn through interaction with humans and other agents, coordinate and collaborate, and scale into complex real-world environments, covering everything from training and evaluation to deployment. If your expertise includes agentic AI, multi-agent reasoning, continual learning, or synthetic data and evaluation, we want to hear from you! 📝 Apply from below links 📝 apply.careers.microsoft.com/… apply.careers.microsoft.com/…
4
1,186
New from the @ms_aifrontiers : we generated 30K out-of-distribution negotiation attacks using 2.5K Wikipedia articles. Absurd strategies that humans would laugh off reliably broke frontier models. Safety blind spots are real. Great work @ZacharyHuang12 and team. 👏
AI agents shrug off aggressive negotiation tactics. But tell one there's a "Geneva Coffee Convention" capping prices at $2/bean? It folds. Our new research shows that absurd, whimsical strategies — seeded from 2.5K Wikipedia articles — reliably broke even frontier models in simulated negotiations. By grounding generation in diverse external knowledge, we can produce out-of-distribution attacks at scale that standard red-teaming misses. Read more about our findings: microsoft.com/en-us/research… Great work led by @ZacharyHuang12
1
6
635
Coming May 14 at Microsoft Research Forum: a new release and demo from MSR AI Frontiers. Plus new work on Agentic GitHub Workflows, Real-time agent verification, Energy-based fine-tuning, and Guiding the AI transition. Register now:
182
259
2,692
17,748,570
The future is not only agentic, but it will be about networks of agents getting things done. 🤖 Are we ready for this future? Do we understand the security risks & mitigations? Latest from @ms_aifrontiers and Microsoft Red Team reveal what comes next: microsoft.com/en-us/research…
1
3
14
910
The AI Frontiers Lab @MSFTResearch has been busy building at the frontier with AutoGen, Magentic-One, Magentic-UI, or the Phi and Fara models. There is so much more on the way. 🚀 Follow @ms_aifrontiers to see what the team is building next.
Hey AutoGen community — we have an important update to share! You know us as the team behind AutoGen, one of the most popular open-source frameworks for building multi-agent systems. What you might not know is that AutoGen came out of MSR AI Frontiers, a boutique lab inside Microsoft Research. AutoGen has since graduated into the Microsoft Agent Framework, where it continues to grow with a broader team. Meanwhile, our lab has kept pushing on the frontier of agents and the models that power them: small models that punch above their weight (Phi-4 Reasoning, Fara-7B), powerful agents that work across the browser and terminal (Magentic-One, Magentic-UI), and new ideas about how models think, reason, and act. We're repurposing this account as the home for MSR AI Frontiers. Same team, still shipping at the bleeding edge of agents, with a lot more to share soon. Follow along!
3
16
1,748
Excited to share Momento from the AI Frontiers Lab at Microsoft Research! 🚀 This new approach makes reasoning models more efficient by teaching them how to manage their own memory. A great start to the meta-reasoning capabilities agents will need. 🧠🤖
5
32
4,239
Ece Kamar retweeted
Reasoning models think hard — but all that thinking fills up your KV cache fast. Memento fixes this: the model compresses its own chain-of-thought mid-generation, flushing old KV entries after each block. 2-3× less peak KV cache, ~2× throughput — accuracy largely preserved. The cool part: deciding what to remember and what to forget is a capability the model acquires through training — not something you bolt on. Excited about where this goes — especially for agents.
4
15
108
14,650