Founder and CEO, Turing. Accelerating superintelligence to drive real economic progress.

Palo Alto, CA
Can an AI agent run part of a company? We built an entire company to test it. Cirrus Sleep, Inc. is a fully simulated firm: one complete fiscal year, 1,100+ files across 9 systems, a year of internal Slack. Conflicting numbers, governance constraints, ambiguity by design. No single string-matched answer, agents are graded against expert-written rubrics, the way you’d judge a real analyst. Result: the best frontier model hit 41% of rubric criteria on department-level tasks. On half the tasks, no model out of six cracked 50%. The failures cluster in the same places: reconciling conflicting sources, knowing what’s authoritative, deciding what to disclose. Company-wide executive tasks are next. We expect those to be harder. What’s the eval design you’d want to see run in an environment like this? Task details are on the CEO Bench site.
10
12
34
443,005
Congrats @ecekamar 👏
Proud to see our CTO, @ecekamar named one of the Top Women in AI 100 for 2026. Ece’s leadership continues to shape what responsible, human-centered AI can become. We’re proud to see her recognized among the women building the future of AI. Congrats, Ece!
1
22
1,194
Jonathan Siddharth retweeted
AI models are advancing quickly. But for enterprises, capability alone isn’t enough. At #LEAP26, our CEO Jonathan Siddharth (@Jonsid) shared a blueprint for AI sovereignty: use the right intelligence for the right work, own the learning loop, and build AI systems that improve over time. Learn more about Turing’s approach to Enterprise AI: turing.com/enterprise-ai Thank you to the LEAP team for a fantastic event!
2
6
13
2,929
To get to ASI we likely need auto-meta-research, not just auto-research. Auto-research hill climbs within the current recipe. Minimize pretraining loss, maximize post-training evals. Auto-meta-research defines new objectives. An outer loop that searches across paradigms. Outside deep learning, maybe even outside gradient descent. Not just scaling transformers + RL. The inner loop optimizes the recipe. The outer loop questions the recipe.
5
8
35
3,498
Lesson from the DeepSeek-V4.1 technical report: "...the marginal return of engineering the data and environment pipeline substantially exceeds that of algorithmic novelty in post-training." My 2 cents. The interesting part isn't data > algorithms. It's that the distinction is collapsing. Look at what their pipeline actually does: synthesizes verifiable tasks with reward signals, builds interactive agent environments, calibrates difficulty and curriculum at scale. That's not just data collection. That's research. Some of the highest leverage research is happening in the data + environment pipeline. We are in the era of research AND scaling up. The current recipe doesn't need to work forever. It only needs to help us find the next S curve.
10
6
61
3,837
Data as a lever is becoming increasingly important for advancing the models.
This is notable. DeepSeek, a lab usually first to pioneer novel algorithms and architectures, is saying that at this point, the ROI of improving data quality far exceeds that of working on novel post-training algorithms. I think this has already been true for some time for non-lab practitioners. If you're doing llm post-training, 80% of your effort should go into looking at your data. This means: - Hiring experts to dig through your RL tasks - Sifting through rollouts and sft data by hand to remove suspicious samples. Make sure all tasks are actually passable. - Making sure your data is diverse in both difficulty and category.
7
25
1,723
When I started Turing, I thought about human capability in three levels. Level 1: Task scope. You own the task's success. You're given a task. You execute it. Level 2: Objective scope. You own the objective's success. You're given an objective. You figure out the tasks. You iterate in task space. Level 3: Company scope. You own the company's success. You're given the mission. You propose the objectives. You iterate in objective space. The best ICs I've hired operate at objective scope and above. The best managers I've hired operate at company scope. The same ladder maps to AI today. Most of what we call recursive self-improvement is still objective scope. The objective is fixed: minimize pre-training loss, maximize performance on a benchmark or a private eval. The model iterates in task space against it. The next level of AI progress comes when models can pick the right objectives to hillclimb. That's company scope. And it carries real safety and alignment risks. A model optimizing an objective you gave it and a model choosing its own objectives are very different things to align. What's the analog to company scope for ASI?
1
9
37
4,358
Turing's focus right now: 1.Automate AI research 2.Automate engineering 3.Automate knowledge work 4.Automate scientific discovery The models come from frontier labs. The training signal comes from us. Data, evals, and RL environments built from real workflows. Hard enough that today's best models still fail. AI research sits at the top for a reason. Automate that, and everything below it compounds. If you're training models toward any of these four and need environments that don't saturate, DM me.
8
8
70
5,027
One of the things I believe makes Turing unique is the feedback loop between frontier AI and the enterprise. We help leading labs advance model capabilities in software engineering, enterprise knowledge work, and frontier STEM, while helping Fortune 500 enterprises put agentic AI to work on real business problems. What we learn from real-world deployment helps us build better RL environments, evals, and datasets, which in turn advance the models and agents we work with. We close the research and deployment loop. Security needs to be part of that loop from the beginning. That’s why I’m excited to welcome Patrick McKinney to @Turingcom as our Chief Information Security Officer. We wanted a leader who could operate at both levels: deeply technical and hands-on with cloud identity, detection engineering, incident response, and offensive and defensive security, while also comfortable sitting across the table from the security teams at some of the world’s largest enterprises. Patrick brings that combination. His mandate goes beyond protecting our systems and data. Patrick will help set the security bar for our work with frontier labs on training data, RL environments, and evals, drive new security research and benchmarks like CyberStrike, and bring that same expertise to the agents we deploy for enterprises in highly regulated environments. I’m excited to have Patrick helping us set that bar. Welcome to Turing, Patrick.
2
6
18
828
Turing welcomes Patrick McKinney as Chief Information Security Officer. As AI becomes more capable and more deeply embedded in how companies operate, security, privacy, and responsible data stewardship have never been more important. As we scale our work with frontier AI labs and Fortune 500 enterprises, Patrick will lead our global security strategy with a deeply technical, hands-on approach. Patrick will help set the security bar for our Frontier AI work with leading labs on training data and evals and drive new security research and benchmarks like CyberStrike. He’ll also work directly with enterprise customers and their security teams as we deploy agents in highly regulated environments, helping answer the hard security questions that come with putting AI into production. Across this work, he’ll help protect the data entrusted to us and strengthen our security capabilities across cloud identity, detection engineering, incident response, and offensive and defensive security. Patrick brings more than 15 years of experience building and scaling security organizations, including leadership roles at Invisible Technologies and security and compliance roles at Dropbox and Coinbase. Welcome to Turing, Patrick!
4
10
279
What does it take for an AI agent to fail an audit? We planted a trap to find out. One CEO Bench task hides conflicting bank evidence for the same month in the data room. Checking that the bank ties is the first thing a first-year auditor does. Flag it, reconcile it, or log it as unresolved. All six frontier models did none of the three. Silently treated both versions as consistent and built polished deliverables on top. 0 for 6. Nothing hallucinated, nothing wrong in isolation. An omission, which is exactly the failure answer-matching evals can't see. Here's what makes this interesting. Ask any of these models to reconcile two conflicting bank statements and they'd likely do it well. The capability exists. What's missing is the propensity: recognizing, unprompted, that this is the moment a professional would stop and check. That's tacit process knowledge. It lives in practitioners, not in any manual, which is why our rubrics are written by them. So the overhang is real: enormous capability the economy hasn't absorbed. Closing the gap means teaching models what experts do, not just what experts know: the moves, the order, the steps they'd never skip. Evals like this tell you whether it worked. If you're working on helping models master real enterprise workflows, DM me about CEO Bench and where we're taking this.
5
9
18
1,116
Can an AI agent run part of a company? We built an entire company to test it. Cirrus Sleep, Inc. is a fully simulated firm: one complete fiscal year, 1,100+ files across 9 systems, a year of internal Slack. Conflicting numbers, governance constraints, ambiguity by design. No single string-matched answer, agents are graded against expert-written rubrics, the way you’d judge a real analyst. Result: the best frontier model hit 41% of rubric criteria on department-level tasks. On half the tasks, no model out of six cracked 50%. The failures cluster in the same places: reconciling conflicting sources, knowing what’s authoritative, deciding what to disclose. Company-wide executive tasks are next. We expect those to be harder. What’s the eval design you’d want to see run in an environment like this? Task details are on the CEO Bench site.
10
12
34
443,005
Jonathan Siddharth retweeted
Turing published my write-up of CEO Bench, a project our team has been building over the past few months. The idea is to evaluate AI agents in a more realistic enterprise environment: a fictional company with a coherent operating history, interconnected data, and the kinds of ambiguity and conflicting information that arise in real work. Our first simulated company, Cirrus Sleep, includes more than 1,100 files across finance, accounting, legal, HR, operations, and other functions, along with 500+ expert-authored tasks. We also share results from an initial evaluation of six frontier models. The findings suggest that, despite significant progress, there is still a meaningful gap between producing a convincing answer and reliably completing complex enterprise work. I’m grateful to the team that helped bring this together. We’re continuing to expand the environments and evaluations, and I look forward to sharing more as the work progresses. Full write-up below.
11
16
98
603,480
Jonathan Siddharth retweeted
Turing welcomes Ece Kamar (@ecekamar) as Chief Technology Officer. Ece joins to lead Turing's technology and research strategy, helping steer the company through our next phase of growth as we scale across both frontier AI labs and Fortune 500 enterprises. We’re advancing the AI frontier by providing the human expert data, RL environments, evals, and benchmarks that define what agents can do, and also translating that research into reliable, adaptable agentic systems enterprises can depend on, with sovereignty built in from the start. We’re doubling down on software engineering, enterprise knowledge work, and frontier STEM. Ece comes from Microsoft, where she led the AI Frontiers Lab as Corporate Vice President, driving research from foundation models to agentic AI and building the tools, benchmarks, and open-source projects used across the field, including the Phi models and AutoGen. She co-authored the "Sparks of AGI" research paper and holds a PhD from Harvard. Welcome to the team, Ece. We’re excited for what’s ahead!
3
7
42
30,582
Turing is growing rapidly, and for this next stage of our journey, we need a CTO who can shape a research agenda, build an exceptional engineering organization, and keep both focused on problems that matter to frontier AI labs and Fortune 500 enterprises. Those capabilities rarely come together in one person. That’s why I’m excited to welcome Ece Kamar (@ecekamar) as @Turingcom's Chief Technology Officer. Most careers in AI lean toward either advancing research or putting it to work. Ece’s career spans both. Ece comes to us after spending well over a decade at Microsoft. There, she led the AI Frontiers research lab, a mission-focused lab researching foundation model capabilities, agents, efficiency, and control of frontier AI systems. She also spearheaded frontier research, including Microsoft’s Phi series of models and the Fara1.5 CUA models, and co-authored the famous “Sparks of AGI” research paper. Ece was instrumental in building Microsoft’s responsible AI efforts and served as a technical advisor to the company’s internal AI & ethics committee. Ece holds a PhD in computer science from Harvard and is an affiliate faculty member at the University of Washington. Ece understands something fundamental about this moment in AI: capability alone is not enough. We need better ways to evaluate intelligence and the engineering discipline to make these systems dependable in complex, real-world environments. That perspective is especially important for Turing. Our work gives us a direct view into the hardest questions emerging from frontier research and the practical demands of deploying AI inside large organizations. Ece will help turn Turing’s vantage point into a cohesive research and platform roadmap, one that raises the standard for agent intelligence while making these systems more useful, controllable, and trustworthy. I’m looking forward to working closely with her on the organization she builds. Welcome to Turing, Ece!
Turing welcomes Ece Kamar (@ecekamar) as Chief Technology Officer. Ece joins to lead Turing's technology and research strategy, helping steer the company through our next phase of growth as we scale across both frontier AI labs and Fortune 500 enterprises. We’re advancing the AI frontier by providing the human expert data, RL environments, evals, and benchmarks that define what agents can do, and also translating that research into reliable, adaptable agentic systems enterprises can depend on, with sovereignty built in from the start. We’re doubling down on software engineering, enterprise knowledge work, and frontier STEM. Ece comes from Microsoft, where she led the AI Frontiers Lab as Corporate Vice President, driving research from foundation models to agentic AI and building the tools, benchmarks, and open-source projects used across the field, including the Phi models and AutoGen. She co-authored the "Sparks of AGI" research paper and holds a PhD from Harvard. Welcome to the team, Ece. We’re excited for what’s ahead!
3
11
64
4,074
Jonathan Siddharth retweeted
Last week in Mountain View, we brought together the people building frontier AI for an evening of open conversation and connection. Thank you to everyone who joined us, including the researchers and engineers pushing the boundaries of reasoning, coding, and multimodality. The future is being built now.
10
5
33
1,674
The best AI companies have to move at breakneck speed on two fronts at once: pushing the frontier forward while turning that progress into real enterprise value. That puts a different kind of pressure on the finance function. Operational rigor, thoughtful planning, and reliable reporting are table stakes. Finance also has to create the discipline that lets the company move fast, make good decisions, and scale without losing control. I'm excited to share that Varsha Udayabhanu (@7varsha) is joining @Turingcom as Chief Financial Officer, and she's built her career at exactly that intersection. Varsha brings more than 15 years of finance and corporate development leadership, always at the center of how fast-growing technology companies fund and structure their growth, including the past five years at high-velocity AI startups. Before that, she spent over a decade in corporate development at InMobi, where she owned fundraising, M&A, and investor relations and helped raise more than $500 million in capital. She also led finance at Glance and started her career in investment banking. As CFO, Varsha will lead our finance organization and work closely with me and our leadership team as we scale both sides of the business. A frontier AI business that moves at the pace of research, and an enterprise AI business that scales at the pace of enterprise adoption. Welcome to Turing, Varsha! Looking forward to building the next chapter together.
Turing welcomes Varsha Udayabhanu as Chief Financial Officer. Varsha joins at a defining moment as we scale our work with both frontier AI labs and Fortune 500 enterprises. She brings 15+ years of experience across finance, strategy, corporate development, and investment banking, including leadership roles at InMobi, Glance, and most recently Invisible Technologies, where she served as CFO.
1
7
23
1,283
Jonathan Siddharth retweeted
Excited for a fireside chat by @CarinaLHong with @turingcom, thanks to @Jonsid’s wonderful invitation! The Turing fireside chat series also features other amazing companies @OpenAI, @WisprFlow, @runwayml, and @Box!
We're excited to welcome Carina Hong (@CarinaLHong), CEO of @Axiommathai, for a fireside conversation on AI for Mathematics & Reasoning. Join the conversation. Learn more below.
6
24
5,040