Glad to announce our Series A, bringing on some incredible partners to support us in solving one of the fundamental problems in AI.
Today we're announcing our $40M Series A at a $400M valuation, led by @a16z , with participation from existing investors @8vc, @pearvc, and @BloombergBeta and new investors @HRTVentures and @nextladder. Alongside the fundraise, three more announcements: - Vals Smith: is now generally available. Anyone can create a custom coding benchmark from any GitHub repo with 120 free credits to get started. - Frontier Risk Benchmarks: We are releasing the RSI Index in collaboration with @CoreWeave and just launched ReverseEngBench, a new cyber benchmark built with Columbia University, Tufts University, UC Berkeley, and UCLA. We are also sharing our initial work in mental health, with more to come across environmental impact, military, and biosecurity. - Website + Vals Index 2.0: We completely rebuilt the Vals website and have released Vals Index 2.0, with coverage of more of the economy. Our revenue has already grown 8x compared to all of 2025. Our customer base doubled and the team tripled in 6 months. Our results have been cited in model cards from OpenAI, Anthropic, Google, Meta, and xAI. The AI economy runs on self-reported grades. When a model ships, the scores come from the company that built it. No other trillion-dollar industry works this way. Finance has ratings agencies. Medicine has the FDA. AI has vibes and vendor benchmarks. We built Vals to be the independent evaluation layer the industry is missing. Check out our new website and try Vals Smith!
51
15
217
58,180
my take, both labs have miscalibrated their medium reasoning levels. opus 5.5 is too expensive for its performance, astra is better than it needs to be
How does "thinking harder" change how models perform on math proofs? We ran GPT-6 Astra, Opus 5, and Opus 5.5 on every reasoning level: low, medium, high, xhigh, and max on our benchmark, Proof Bench v1.1.
1
8
1,511
forget the hotline its 2026 we need a cross org slack channel
3
11
932
Rayan Krishnan in DC retweeted
The Washington Post just released an article about our work on AI Child Safety. They key takeaway: child safety can’t be measured from a chatbot’s first answer alone. Across nine models and 648 simulated, 10-turn teen conversations, at least one critical safety check failed in 27.5%. Of those conversations, 62% included a failure later in the exchange.
3
5
17
2,389
I'm in DC to kick off our public sector summit. Excited to host folks from AI labs, Congress, embassies, and think tanks. I'm hosting a panel on AI evaluation with FRONTIER Act staffers. There will also be sessions on US China AI policy, child safety and others. We have a few spots left. DM me if you'd like to join or connect while I'm in town.
2
2
22
987
We’ve updated our timeline for full RSI to July 2027 instead of August after our eval of Opus 5.5. It’s telling that the biggest advocate for pacing is still very much racing ahead. Even the most rational can fall prey to this multi polar trap. The only solution is coordination.
Claude Opus 5.5 takes #1 on RSI Index and is the first model to beat the published reference on LM Training under our protocol, marking a major step forward for long-horizon agentic work.
30
79
1,093
86,920
People vastly underestimate the capability + risks unlocked by multi-agent swarms. Simply put, Einstein could not single-handedly build a rocket. The greatest feats of human intelligence are a consequence of societies of well-organized humans or even civilizations with humans working over generations. We already see glimpses of the step-level changes in the ability unlocked by coordinated agents. Navier Stokes is a great example of this where we got essentially 100 years worth of single agent work compressed into 88 hours by 10k agents. The parallel evaluation challenge is significant: 1) we need more thorough containment strategies (to prevent OAI-HF type issues) 2) the alignment problem becomes two layer (alignment of the agent to the swarm and the swarm to the user intent) 3) an incredible volume of metadata is produced making it harder to monitor / investigate what is actually happening.
2
4
29
2,612
Rayan Krishnan in DC retweeted
We're excited to announce our AI Engineer New York speaker lineup! Presented by @arizeai CEO @ExaAILabs CTO @tryramp CEO @ValsAI CEO @RogoAI CTO @modal CEO @trybasis CEO @SonarSource CEO @runwayml CTO @mintlify CEO @sfcompute CTO Bridgewater CTO @turbopuffer CEO @NousResearch CEO @gen_intuition HEAD OF AI @coatuemgmt TECH LEAD @Vanguard_Group MTS @GoogleDeepMind HEAD OF DEVX @cerebras HEAD OF US BD @MiniMax_AI DEPUTY CISO @Mastercard COO @citsecurities SR. STAFF ENGINEER @GEICO HEAD OF ENGINEERING @linear DIRECTOR OF DX @togethercompute HEAD OF AI PRODUCT @coinbase AI INVESTMENT LEAD @BlackRock GROUP HEAD OF DATA & AI @Wise HEAD OF AI @FTI_Global KNOWLEDGE GRAPHS LEAD @Point72Careers DIRECTOR, AI & DATA @AlixPartnersLLP ANALYTICS ENGINEERING LEADS @monzo DISTINGUISHED ENGINEER @CapitalOne OPS & INFRASTRUCTURE LEADER @RevolutApp GLOBAL CHIEF SECURITY ARCHITECT @jpmorgan CORPORATE VP @NewYorkLife HEAD OF FRONTIER AI MODELS @WellsFargo AGENT HARNESS ENGINEER @cursor_ai / @xai COFOUNDER @getcompoundai (ex-VP @coatuemgmt + @ycombinator) DIRECTOR OF DATA SCIENCE @Fidelity HEAD OF DATA, DIGITAL & AI @apolloglobal SVP, ENTERPRISE PLATFORM ENGINEERING ARCHITECTURE @twosigma PARTNER, HEAD OF THEMATIC INVESTING @apolloglobal HEAD OF ARCHITECTURE & EMERGING TECH, OFFICE OF THE CTO @Bloomberg PRODUCT + ENGINEERING LEAD, CHATGPT FOR FINANCIAL SERVICES @OpenAI PRODUCT + ENGINEERING LEADS, CLAUDE FOR FINANCIAL SERVICES @AnthropicAI Oct 12–14 in NYC. ai.engineer/nyc
22
19
327
35,705
astra inference time scaling is so clean. at varying reasoning levels its capable of consistently greater intelligence without wasting tokens. i've been finding the model to be fast / super token efficient so ran this mini experiment. i suspect this is the underlying reason why the model can take on challenging tasks and games without being painstakingly slow. not sure if its scaling rl, looping, rsi or witchcraft but im impressed.
2
1
84
4,798
Rayan Krishnan in DC retweeted
This is why having standardized benchmarking methodology across models is important!
Epoch AI researcher Michelle Campeau explains how a model can look 40% better overnight simply because the benchmark changed, not the model: "Being able to see this is the highest score on this specific benchmark is the best information you have about how something is going to do." "Critical Point, one of the physics benchmarks, Anthropic mentioned in their latest system card that it has over 40% of questions that are broken, so they ran their own corrected version. And this corrected version is private to them, so no one else can understand how these scores compare to past scores." "You might compare anyways and be like, whoa, the scores are 40% better than they used to be. And actually, that is not the takeaway you should have." @EpochAIResearch
1
3
773
Rayan Krishnan in DC retweeted
AI models are advancing faster than legacy benchmarks can keep up. @TechCrunch @LucasRopek1 visited us to see how we’re building independent evaluations grounded in real-world work and designed to measure both capability and risk
Vals AI is hoping to make AI benchmarking a more neutral and trustworthy resource in a world increasingly inundated by AI models. spr.ly/6016BGku1M
5
1
42
6,636
Independent evaluation requires meaningful access, real independence, transparency, and the freedom to publish findings. I signed the AEF statement because these principles closely align with how we approach evaluation at Vals.
Today, more than 100 leading AI experts endorsed a set of minimum requirements to take seriously AI companies' recent call to embed external evaluators. These evaluators need to be genuinely independent, transparent, and represent a range of expertise areas. They also need to be guaranteed employee-level access and to be protected from retaliation for findings that make companies look bad. We welcome model developers’ recent calls for independent oversight, but it’s what they do next that matters. The labs must be accountable for ensuring these requirements are met, so that the public can have faith in the process and the outcomes. Over the past week, the AI community has debated the appropriate role of external evaluation, including who should do it and on what terms. We may not agree on everything, but there is a lot of common ground. To make embedded evaluations credible, more than 100 experts with varying backgrounds and ideas about AI risk agree in today’s letter that frontier AI developers should: 1. Guarantee embedded evaluators full editorial independence and mitigate conflicts of interest 2. Rely on multiple evaluators with differing viewpoints and areas of expertise 3. Publicly document the terms under which evaluators operate, as well as facilitating permissive publication of methods and findings 4. Shield evaluators from retaliation 5. Grant access equivalent to that of highly privileged employees There is a thriving and growing ecosystem of independent AI evaluators who are advancing this science every day – but we need aligned standards, guaranteed protections, and independent funding. That’s why we created the AI Evaluator Forum. Today we are entering our next phase. We’re launching an open call for new members, collaborators, and independent funding sources to help evaluators meet this moment and demand accountability from developers. Join us in building the evaluator ecosystem. See the public letter here: aievaluatorforum.org/initiat… Learn more at aievaluatorforum.org/path-ah…
2
23
2,061
Rayan Krishnan in DC retweeted
Astra for Law shows a HUGE performance increase in our Legal Research Bench
Astra for Law: Frontier intelligence built for your practice. A new offering powered by GPT-6 Astra with tools, settings, and context to support the expertise and judgment of lawyers and legal technology firms.
3
19
455
119,773
Rayan Krishnan in DC retweeted
Existing coding benchmarks stop at the first working version. We are releasing Vibe Code Bench 1-100 today, to measure what comes next. This benchmark asks if models can handle a large number of modifications to a product, without breaking what already works.
21
8
128
17,553
People underestimate how much better Astra is than Sol in its ability to have novel discoveries. Would not expect this trend to slow.
Scientific discovery is the next frontier for AI systems. However, new scientific results are difficult to verify, and therefore hard to measure. Today we're releasing MysteryMechanism, a benchmark that tests whether agents can rediscover sealed mathematical mechanisms through bounded experiments. It cuts sharply across the frontier, with Astra landing roughly 20pp above Sol.
2
13
1,142
Rayan Krishnan in DC retweeted
astra is the first ai that i believe: + given enough time it would solve anything. even after the fateful creeper explosion of all valuable stuff, it can learn from the lesson and overcome with this message: "The new chest has been destroyed, and the stored valuables are missing. I’ll keep all future critical items in my inventory, where keepInventory protects them, and rebuild the supplies while continuing toward the dragon." you can also see in the chart there are basically no plateau whatsoever. + has emergent behaviors that are unusual to humans, while still showing some familiar traits. in its messages, astra has show some emotions frustration or self-criticism, but at the same time it counts pixel when aiming a bow and writes its scratchpad in increasing more cryptic ways. i believe that if we can see raw chain of thoughts, there might be more evidence of its intentions and emotions behind its actions. ai's thoughts and behaviors may diverge from ours and become completely strange to humans, and we would need to put in much more effort just to understand it. watching astra is more more similar to watching an exotic living creature now. it's strange
GPT-6 Astra had gotten further than any AI system had ever gone in Minecraft. It was able to set up a semi-automatic blaze farm, allowing it to collect 6 blaze rods. It then located a warped forest, where it killed 6+ endermen and collected 3 pearls. As thousands of viewers watched the stream, Astra put all of its valuable items in a chest. Then a creeper showed up and blew up both the chest and Astra’s bed, wiping out all of its progress. Astra seemed to get extremely frustrated after this happened. Astra said to itself, "ALWAYS CARRY CRITICALITEMS withkeepInventory; don'tstoreinunguardedchesteveragain."
12
9
326
28,014
Rayan Krishnan in DC retweeted
GPT-6 Astra had gotten further than any AI system had ever gone in Minecraft. It was able to set up a semi-automatic blaze farm, allowing it to collect 6 blaze rods. It then located a warped forest, where it killed 6+ endermen and collected 3 pearls. As thousands of viewers watched the stream, Astra put all of its valuable items in a chest. Then a creeper showed up and blew up both the chest and Astra’s bed, wiping out all of its progress. Astra seemed to get extremely frustrated after this happened. Astra said to itself, "ALWAYS CARRY CRITICALITEMS withkeepInventory; don'tstoreinunguardedchesteveragain."
129
497
7,937
972,949
Joined @EdLudlow to talk about AI progress, independent evaluation, and why I’m optimistic. Some quick takes: - We’re better at building AI than understanding it. Attention towards testing/evaluation matters more than slowing down. - Our RSI Index projects models could match human researchers on the tasks we test by August 2027. Embedded evaluators can produce more accurate estimates based on internal systems. - Public conflict masks cooperation. The labs, policymakers, and enterprises we work with want better evidence. I’ve seen enough to believe coordination is possible. - Independence comes at a cost. We’ve rejected contracts that would compromise ours. The same group doing the testing shouldn’t also sell the solution. - Evaluation should scale through better technology. If it becomes a bureaucratic moat for incumbent labs, we’ve failed. - Market-based evaluation has a role with or without regulation. Competition pushes us to build better technology and keep up with the frontier.
8
14
68
16,474
Rayan Krishnan in DC retweeted
"My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic..." - @DarioAmodei in his weekend essay. We are going to look very closely at this idea on today's show with @RayanKrishnan. Here is @ValsAI RSI Index. The chart shows rapid improvement in models’ ability to autonomously do AI R&D over roughly the past year. From Vals: Can a model do the research that builds the next model? Autonomous LLM R&D scored against selected published, human, and frontier-model references. Or simply: can an AI model act like an AI researcher — running experiments and improving AI systems — with little or no human help? Vals' conclusion seems to be: AI is getting good at doing AI research — but not yet at inventing the breakthroughs that would drive recursive self-improvement.
3
2
24
8,235
With AI safety topics going mainstream, the public seems anxious and disconnected from the reality of the problem. From what I’ve seen at Vals AI, I’m optimistic we’ll coordinate toward an optimal future for AI. Our study on RSI shows that, at their current pace, Anthropic’s models could match human researchers by August 2027. That creates urgency but gives us time to prepare. It’s hard for the public to know whom to trust when everyone debating has their own incentives. This was the concern I had when I started Vals AI: that a multipolar paradox would emerge, where the actions of self-interested parties lead to a non-optimal outcome for the system. To overcome this, we independently evaluate models for their real-world impact. This mirrors the role of auditing firms. I’m optimistic because we have found rational and willing partners across the industry. Every major lab has been a great collaborator, providing us with early access to models for testing on our public benchmarks. Every member of Congress and government agency we’ve briefed has been eager to learn. Enterprises are becoming more sophisticated about adopting models based on evaluated capabilities/risks. It hasn't been easy. We have had to earn the trust of competing groups and reject significant contracts that would have compromised our independence. But done right, evaluation can scale with the frontier through automated infrastructure. Embedded evaluators can understand systems during development while maintaining independence. This makes it easier for new entrants to compete and promotes transparency that builds trust with the public. We’re eager to see a diverse ecosystem emerge. We’ve open-sourced our core infrastructure and published our methods, supporting peers. This is a time of high variance. The decisions we make now will have an outsized impact on the future we arrive at. I remain optimistic we will get this right in the year ahead.
7
17
81
15,327
Rayan Krishnan in DC retweeted
Every few months in 2026: 「Hey guys, we should really consider slowing down. Our models are getting dangerous.」 CEOs: I agree. 1 month later 「Our next frontier model accidentally hacked into NASA and solved P=NP. Still in training.」
1
1
9
668