Pinned Tweet
Proximal builds infrastructure that enables AI to improve from real-world experience. In the past year, we have grown to more than $200M in annualized revenue helping frontier labs and enterprises improve coding agents. Now, we are expanding beyond software engineering. Recent progress in mathematics demonstrates what happens when models are trained in domains with massive amounts of public data and easy verifiability, making it easy to find weaknesses, generate training tasks, and iterate quickly. We build infrastructure that enables this feedback loop in other domains: we are excited about a world in which AI systems design targeted drugs, accelerate chip design, and rewrite legacy software that still runs hospitals, power grids, governments, and other critical infrastructure. Since starting, we have assembled a small team of researchers and engineers and built world-class infrastructure for post-training and synthetic data research to overcome the challenges of scaling manual data curation. Our team comes from Cursor, Google DeepMind, Meta Superintelligence, Prime Intellect, Citadel, and Jane Street. More than half are former founders. We’re fortunate to be backed by investors who share that vision, including @generalcatalyst who led our $15M Seed Round at a $300M valuation, as well as @chemistry, @svangel, @dvlamoen and individuals like @LiamFedus, @kevinweil, and @bernhardsson
112
41
371
233,927
Proximal retweeted
If you are interested in: - Post-Training - Data Research - Infra at tremendous scale or just generally want to know more about what we do @ProximalHQ, please reach out, I would love to chat :)
Excited to share i’ve joined @ProximalHQ here in SF! I think that there is a ton of interesting open problems related to data and post training at large. There is no better team to work with than the one we have and I am extremely excited to share our work :)
15
15
470
21,939
Many cybersecurity evals use differential execution to test whether agents can reproduce known software vulnerabilities Without added constraints, this creates fairness issues: Our evaluation shows that CyberGym is effectively saturated on a verified task subset
10
11
75
17,791
For more reliable cybersecurity evals, most robust graders are required: SEC-Bench Pro achieves high precision/recall on vuln-repro tasks with LLM-as-judge, and ExploitGym additionally uses a Capture-the-Flag approach for clear certainty We consider these benchmarks high quality and recommend using them for cybersecurity evaluations
1
7
392
Read our research blog here: proximal.ai/blog/cybergym-cy…
5
348
Gemini 4 Argon scores 55.0% on FrontierSWE The model is close to Fable 5.1 and only outperformed by Opus 5.5 and GPT-6 Astra
Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.
4
13
209
16,110
Compared to previous versions of Gemini, Argon works for much longer. Despite this, it is slightly more time-efficient than other top-scoring models, placing it on the pareto frontier
1
11
717
The models that will accelerate drug discovery, chip design, and energy are limited by the data they learn from. We led @ProximalHQ’s $15M seed as they build the data engine for the next generation of frontier AI and treat training data with the same rigor as model architecture. Watch @quentinclark in conversation with @calvinchen. Chapters 00:00 — Meet Proximal 03:45 — A Rainy Afternoon in South Park 05:19 — From Sneaker Bots to a First Exit 08:14 — Why Suffering Equals Growth 10:13 — The Path to Proximal 11:57 — How Models Learn 15:17 — Training Data as a Research Problem 23:24 — Building the Team in India
10
6
52
13,369
Proximal builds infrastructure that enables AI to improve from real-world experience. In the past year, we have grown to more than $200M in annualized revenue helping frontier labs and enterprises improve coding agents. Now, we are expanding beyond software engineering. Recent progress in mathematics demonstrates what happens when models are trained in domains with massive amounts of public data and easy verifiability, making it easy to find weaknesses, generate training tasks, and iterate quickly. We build infrastructure that enables this feedback loop in other domains: we are excited about a world in which AI systems design targeted drugs, accelerate chip design, and rewrite legacy software that still runs hospitals, power grids, governments, and other critical infrastructure. Since starting, we have assembled a small team of researchers and engineers and built world-class infrastructure for post-training and synthetic data research to overcome the challenges of scaling manual data curation. Our team comes from Cursor, Google DeepMind, Meta Superintelligence, Prime Intellect, Citadel, and Jane Street. More than half are former founders. We’re fortunate to be backed by investors who share that vision, including @generalcatalyst who led our $15M Seed Round at a $300M valuation, as well as @chemistry, @svangel, @dvlamoen and individuals like @LiamFedus, @kevinweil, and @bernhardsson
112
41
371
233,927
Claude Opus 5.5 ranks #2 on FrontierSWE The model achieves a score of 62.3%, closely behind GPT-6 Astra (65.5%) and clearly surpassing Fable 5.1 (56.3%) and Opus 5 (52.0%)
8
6
167
12,699
We see large improvements in tasks that require vision capabilities - Opus strongly improves on TORCS Racing Bot, Flight-Sim Renderer in OpenGL and Fitness Recap Video in Remotion. For more information on Opus 5.5 on FrontierSWE v2, check out the model card
1
9
783
We partnered with @FireworksAI_HQ to bring FrontierSWE v2 to the Specialized Intelligence Index Evals are essential for improving and safely deploying AI. We're excited to support the initiative!
Today we're launching the Specialized Intelligence Index (SII): one destination for real-work benchmarks across industries, built by the teams that use them every day. Hear from Fireworks co-founder @the_bunny_chen on the importance of specialized benchmarks:
1
6
35
3,925