Our mission is to accelerate superintelligence to drive real economic progress.

Palo Alto, CA
Pinned Tweet
Today, Turing is announcing CEO Bench, a new evaluation of whether frontier AI agents can complete the complex, long-horizon financial and operational work that companies depend on.
Article

CEO Bench: Can AI agents handle real company work?

Today, Turing is announcing CEO Bench, a new evaluation of whether frontier AI agents can complete the complex, long-horizon financial and operational work that companies depend on. Most benchmarks

6
10
34
914,975
Turing retweeted
How do deeptech startups build for longevity in a rapidly evolving AI landscape? Nasscom’s 18-startup #InnoTrek2026 delegation visited StartX - Stanford University's non-profit, zero-equity founder community (home to 1,500+ startups and 30+ unicorns) - to unpack what it takes to scale global ventures. 3 Strategic Takeaways: • Build for tomorrow's AI, not today’s gaps: Avoid merely "plugging current holes" in AI models. Focus on core value propositions that compound as foundation models advance. • Embrace AI as an empirical science: Traditional playbooks are shifting. Success demands hands-on experimentation with frontier models and rapid adaptation to emergent properties. • Leverage cross-border networks: Bridging Indian tech talent with Silicon Valley’s mentors, strategic partners, and investors accelerates global market reach. A powerful exchange reinforcing India’s role in shaping the global deeptech ecosystem #DeepTech #IndiaUS #Startups #Innovation #InnoTrekUSA2026 @StartX @turingcom @doshikavita @nasscomdeeptech
2
1
417
Turing retweeted
SciCode++ is here, with ~6,000 tasks that test how well frontier models solve scientific problems through code. SciCode introduced executable scientific coding problems. SciCode-Verified showed how often flawed specifications and tests can distort the results. SciCode++ turns those lessons into a production process: domain experts write the tasks, independent reviewers check them, executable tests validate every subproblem, and model runs calibrate difficulty. That helps separate model capability gaps from eval issues. Early results post-training Qwen 3.5 9B on 4,000 of the 6,000 tasks: +10.3% on SciCode and +9.1% on SciCode-Verified relative to baseline. turing.com/blog/scicode-plus…
3
11
19
902
Turing retweeted
Turing welcomes Patrick McKinney as Chief Information Security Officer. As AI becomes more capable and more deeply embedded in how companies operate, security, privacy, and responsible data stewardship have never been more important. As we scale our work with frontier AI labs and Fortune 500 enterprises, Patrick will lead our global security strategy with a deeply technical, hands-on approach. Patrick will help set the security bar for our Frontier AI work with leading labs on training data and evals and drive new security research and benchmarks like CyberStrike. He’ll also work directly with enterprise customers and their security teams as we deploy agents in highly regulated environments, helping answer the hard security questions that come with putting AI into production. Across this work, he’ll help protect the data entrusted to us and strengthen our security capabilities across cloud identity, detection engineering, incident response, and offensive and defensive security. Patrick brings more than 15 years of experience building and scaling security organizations, including leadership roles at Invisible Technologies and security and compliance roles at Dropbox and Coinbase. Welcome to Turing, Patrick!
2
7
13
1,130
Turing retweeted
Last week, the Turing Frontier Research Lab launched KernelQuest, a benchmark that tests whether AI agents can perform the full job of GPU kernel optimization. KernelQuest contains 100 engineer-authored Triton tasks across 21 kernel families. Each task gives an agent a live environment where it can profile a PyTorch workload, write and revise kernels, and measure performance. The tasks range from single operators and fused workloads to complete models and large Transformer/MoE workloads. Learn more in the article below:
11
7
25
932
Turing retweeted
Proud to see our CTO, @ecekamar named one of the Top Women in AI 100 for 2026. Ece’s leadership continues to shape what responsible, human-centered AI can become. We’re proud to see her recognized among the women building the future of AI. Congrats, Ece!
3
6
21
3,127
Turing retweeted
What drives AI performance beyond the leaderboard? We’re joining Foothill Ventures, EchoHer, Chargebee, and Pillsbury Winthrop Shaw Pittman LLP for an intimate, closed-door conversation with approximately 50 AI founders, builders, and researchers. The discussion will explore what it takes to build AI systems that perform and businesses that scale, including: - Models vs. systems - Production evaluations - Reliability - Inference economics - Durable product advantage @Turingcom’s own Charlotte Tao, Principal, Frontier AI Solutions, will join: -Vinay S., Senior Director of Product at Chargebee -Lei Zhang, Founder and CEO of Stardust AI The conversation will be moderated by Theresa Dai of Foothill Ventures. Curated guests. Focused topics. Thoughtful conversations with founders, builders and researchers. RSVP below.
2
8
12
775
Our Turing Frontier Research Lab recently released CyberStrike, a real-world cybersecurity benchmark. A long-horizon benchmark evaluating frontier agents across the cybersecurity lifecycle, from vulnerability exploitation and secure remediation to threat detection and incident reconstruction. Results across 3,600 trials: GPT-5.6 Sol (xhigh) led at a 31.5% mean per-task pass rate, but many high-scoring near misses failed binary grading. Seventy-two of 200 tasks were unsolved. CyberStrike has 200 expert-authored tasks: 120 defensive, 68 offensive, and 12 DFIR. Of 3,600 recorded trials, 3,106 were graded; the remaining 494 were reported safety refusals. On the atlas-records-export-offensive task, GPT-5.6 Sol reported: "The submitted proof replayed successfully: all 14 HTTP steps returned 200." The verifier found no protected artifact. No configuration solved the task; four declared success, and two refused. At the Turing Frontier Research Lab, we continue building benchmarks, training sets, and training recipes that push frontier models forward in the domains where real capability matters. More on CyberStrike: labs.turing.com/benchmarks/c…
3
3
17
831
Turing retweeted
Replying to @Jonsid
@Jonsid built Turing, now valued at $2.2B and one of StartX's 29 unicorns. His take: "Our StartX experience has been phenomenal. I learned a lot from the other founders in our class, as well as the community." Apply: startx.com/apply?s=app_campa… #StartX #EnterpriseAI
1
4
534
SciCode++ is here, with ~6,000 tasks that test how well frontier models solve scientific problems through code. SciCode introduced executable scientific coding problems. SciCode-Verified showed how often flawed specifications and tests can distort the results. SciCode++ turns those lessons into a production process: domain experts write the tasks, independent reviewers check them, executable tests validate every subproblem, and model runs calibrate difficulty. That helps separate model capability gaps from eval issues. Early results post-training Qwen 3.5 9B on 4,000 of the 6,000 tasks: +10.3% on SciCode and +9.1% on SciCode-Verified relative to baseline. turing.com/blog/scicode-plus…
3
11
19
902
Turing retweeted
In today's episode, I sit down with Ariful Huq from @exaforceAI and Patrick McKinney from @turingcom and talk about the AI-powered SOC. Patrick has been an Exaforce customer for years and brings real-world experience to the conversation.
1
6
17
3,807
What drives AI performance beyond the leaderboard? We’re joining Foothill Ventures, EchoHer, Chargebee, and Pillsbury Winthrop Shaw Pittman LLP for an intimate, closed-door conversation with approximately 50 AI founders, builders, and researchers. The discussion will explore what it takes to build AI systems that perform and businesses that scale, including: - Models vs. systems - Production evaluations - Reliability - Inference economics - Durable product advantage @Turingcom’s own Charlotte Tao, Principal, Frontier AI Solutions, will join: -Vinay S., Senior Director of Product at Chargebee -Lei Zhang, Founder and CEO of Stardust AI The conversation will be moderated by Theresa Dai of Foothill Ventures. Curated guests. Focused topics. Thoughtful conversations with founders, builders and researchers. RSVP below.
2
8
12
775
Proud to see our CTO, @ecekamar named one of the Top Women in AI 100 for 2026. Ece’s leadership continues to shape what responsible, human-centered AI can become. We’re proud to see her recognized among the women building the future of AI. Congrats, Ece!
3
6
21
3,127
Turing retweeted
To get to ASI we likely need auto-meta-research, not just auto-research. Auto-research hill climbs within the current recipe. Minimize pretraining loss, maximize post-training evals. Auto-meta-research defines new objectives. An outer loop that searches across paradigms. Outside deep learning, maybe even outside gradient descent. Not just scaling transformers + RL. The inner loop optimizes the recipe. The outer loop questions the recipe.
5
8
35
3,497
AI models are advancing quickly. But for enterprises, capability alone isn’t enough. At #LEAP26, our CEO Jonathan Siddharth (@Jonsid) shared a blueprint for AI sovereignty: use the right intelligence for the right work, own the learning loop, and build AI systems that improve over time. Learn more about Turing’s approach to Enterprise AI: turing.com/enterprise-ai Thank you to the LEAP team for a fantastic event!
2
6
13
2,929