part-time researcher | scout @a16z | @ycombinator alum

San Francisco, CA
post AGI classics to read during the weekend: - Scenarios for the transition to AGI @akorinek - The Intelligence Curse by @luke_drago_ - AI Enabled Coups by @TomDavidsonX - The Ambition Singularity by @danfaggella - Gradual Disempowerment by @jankulveit - How long before Super-intelligence (1997) - The British Industrial Revolution in Global Perspective (2009) - Situational Awareness The Decade Ahead @leopoldasch - AI Monotheism vs AI Polytheism by @BerenMillidge
2
2
33
5,266
Joan Cabezas retweeted
6-12 months ago the large labs started buying bio data this spend is starting to ramp very aggressively. the labs are somewhat indiscriminate buyers with enormous budgets. the therapeutics market has long had buyers of data (large and small pharma) but they've historically been the exact opposite of huge-budget and indiscriminate this shift will distort the bio startup market in a bunch of ways. there will be a lot of short-termist behavior to try to get in front of this capital firehouse. but unlike the AI data labelers (mercor, et al) and the robotics labelers (mecka, et al), the bio startups serving up this data have a chance to emerge from this period with businesses that don't rely on selling data the best bio startups will use these sudden resources the labs are dumping on them to supercharge their novel assay and work towards independently powerful models and eventually directly produce therapies bio is the final frontier and its a very exciting time ahead for bio startups IMO
27
26
568
55,903
Joan Cabezas retweeted
We investigated a seemingly simple question: how do you grade if AI generated parts or assemblies are correct? Here's how we improved on this with our verifier, FrontierCAD.
9
14
327
76,748
Joan Cabezas retweeted
Backprop has been the only credit assignment algorithm capable of training large neural nets. Introducing Dust, the first zeroth-order method to pretrain transformers to approach and sometimes even *exceed* backprop with large amounts of computation (bitter lesson!) - Dust is on the order of 1,000 to 10,000x more compute efficient than EGGROLL, the state-of-the-art ES method, for training transformers. - We chose pretraining transformers because it's arguably the hardest task possible for zeroth-order optimization. - Search in high dimensions is really misunderstood: zeroth-order usually gets better with overparameterization, not worse, i.e. at a fixed population size, larger models (up to 120x) reach a lower loss than smaller ones. - Backprop-like gradients emerge from a large population of activation perturbations, and the alignment holds up at every scale we tested, up to 1B tokens, which is critical for scaling. The core idea is a *virtual population*, which helps scale to large populations in transformers by perturbing activations in parallel. w/ @bishmdl76, @cs_serdar, @akshayvegesna
55
270
2,355
200,169
Thought Jev is fast? We evaluated Jev on cuaspeedrun.com for computer use agents and found: 1. 1.5x slower than Opus-5.5-xHigh 2. 10x lower accuracy 3. 5x more steps That puts it in the GPT-4 league of performance and Astra-xHigh on time What and Why?? 🤯 🧵
5
9
53
12,133
Joan Cabezas retweeted
Building AI inside your company? These are your people. Two @bryel_labs events at SF @Techweek_ next week for founders, CTOs, and engineering leaders: Tuesday evening: Join me and @vinitra_s from @ScholeAI to discuss what it takes to level up your intelligence within your company. (Co-hosted by @thehousefund) 🗓️: partiful.com/e/wmdrrEPxrjo20… Thursday afternoon: Curious about training your own models? Building your intelligence stack? Come swap notes with other pioneers making the same decisions. 🗓️: partiful.com/e/BgX5jT0yxM3fa…
2
5
19
244,664
Joan Cabezas retweeted
Today, I am excited to announce Animation Bench. Our first benchmark to evaluate frontier models on web animation reconstruction. Frontend generation is a common commercial use case for coding agents. In addition to static layout it contains transitions and interactions triggered by dynamic behaviours. We evaluated four frontier models on 48 tasks from real websites across three axes: 1. Visual Similarity 2. Motion Consistency 3. Layout Correctness. We find that motion consistency is often incomplete or missing.
Introducing Animation Bench, our first benchmark to evaluate frontier models on web animation reconstruction. Coding agents can already recreate visually plausible web animations, this work evaluates if they can reconstruct behavior of a web page.
22
16
256
20,391
Joan Cabezas retweeted
Introducing Synthetic Hospital: an open, fully synthetic longitudinal EHR benchmark with verifiable ground truth! 1,268 patients, 5,602 encounters, zero PHI. Physicians could not reliably distinguish its charts from real ones. 📄 arxiv.org/abs/2609.30027 💻 github.com/sparkcpark/synthe… ✍️ sparkcpark.github.io/posts/f…
54
146
1,305
299,082
Joan Cabezas retweeted
In 12 weeks, we built a research facility that is run entirely by AI. AI designs, executes, and observes experiments end-to-end across biology, chemistry, and materials science. We’re introducing SciUniverse: a benchmark that measures AI’s ability to do real-world scientific research.
241
397
2,791
790,068
Joan Cabezas retweeted
There’s a good chance your open source model is costing more than the frontier. Cheap tokens ≠ cheap tasks. Here, we introduce Intelligence Density, and Density Aware Training, our post-training technique to achieve less wasted compute, better learning, all with no knobs to tune. Enabled by default in every Trajectory model.
23
24
287
56,431
Joan Cabezas retweeted
More people need to know about this lab. qlabs.sh/
5
4
106
15,758
Joan Cabezas retweeted
My new report with @luke__emberson.
AI is getting cheaper more quickly than any other transformative tech in history. At a given level of performance, cost has fallen ~47%/quarter since 2023. That’s 4× faster than DNA sequencing, 6× faster than compute, 18× faster than lithium batteries, and (up to 1973) 54× faster than electricity.
2
8
53
3,414
Joan Cabezas retweeted
Reinforcement learning is having a moment. For the past few months, we’ve been working with @tryterac to put together something special: Convergence — an event for the people building the future of RL. Speakers from @GoogleDeepMind, @databricks, @mercor, @arena, @joinHandshake, @turingcom, @deeptuneai, @contra, @Stanford, @Princeton, @UCBerkeley, and more. Our goal: make this THE event for the RL community. Oct 29 in SF. Come join us → terac.com/events/convergence…
13
8
77
21,235
Joan Cabezas retweeted
Trajectory is at @HackTheNorth, come say hi!! We’ve got some cool merch in store for whoever can come up with the wackiest benchmark ideas
6
3
64
3,187
Joan Cabezas retweeted
Introducing WeirdML v3, a fully agentic benchmark featuring 11 complex hand-made tasks. Models must explore and understand unfamiliar data, develop ML and data analysis pipelines and produce results despite limited data, unspecified goals and/or very limited feedback. 1/8
WeirdML v2 is now out! The update includes a bunch of new tasks (now 19 tasks total, up from 6), and results from all the latest models. We now also track api costs and other metadata which give more insight into the different models. The new results are shown in these two figures. The first one shows an overview of the overall results as well as the results on individual tasks, in addition to various metadata. The second figure shows cost vs performance and shows a clear scaling with better results for higher costs. We also have a very varied pareto frontier with 11 models from 6 different companies having the best accuracy for a given cost for at least some of the cost range. Grok 3, Claude Opus 4 and GPT 4.5 are the ones that underperform for their costs, while Gemini pro and o3 pro have the best results at the highest costs. Qwen3 30B3A, grok 3 mini and deepseek R1 also each represent a good chunk of the pareto frontier.
54
56
655
140,614
Joan Cabezas retweeted
We have something very special to share today! We developed a standard to assess benchmark quality, which we plan to apply to the most-used AI capability metrics going forward. We hope this will raise the bar for designing and interpreting evaluations. Let us know what you think!
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
14
8
130
5,407
Joan Cabezas retweeted
today, the intersection of brilliant hardware engineers who religiously use AI is tiny. we believe the tam of new hw engineers will 100x over the next years. If you deeply care about creating an abundance of hardware, this is your bat signal! forms.gle/obkVU66docAuX1xx7
I made a list of great startups to join. It's called the Breakout List. The list has 92 companies. These are the 20 with 25 or fewer employees: - Hone (@moritz_stephan, @CarloWillem, @oqbrady) - Normal (@ansonyuu, @hudzah) - Standard Intelligence (@G413N, @devanshpandey) - Tacit Labs (@ninklefitz, @AmDroste) - American Terawatt (@atroyn, @rslparker, @aranibatta) - Conduit (@clemvonstengel, @riopopper) - Convergent (Omkar Savant, Vivek Katara, @debnilsur) - Core Automation (@MillionInt, @_arohan_) - Engram (@dan_biderman, @EyubogluSabri, @realJessyLin) - Instinct (@noahrshinn) - Keenable (@styskin, Matthias Petri) - Lumaril (Mark Elliot, Ben Duffield) - Neion Bio (@Dimkell, Sam Levin) - Pangram Labs (@max_spero_, @bradley_emi) - Quadrillion (@echinaceous) - Re (@karnsaroya, @AnandDhillon, @thecliffwhite, @benaneesh) - Ricursive (@annadgoldie, @Azaliamirh) - Sail Research (@neilmovva, @blintzbase) - Trajectory (@rronak_, @michaelelabd, @QuantumArjun) - Watney Robotics (Sean Cheong, Ryan Gannon) Picks from Elad Gil, Charlie Songhurst, Keith Rabois, Mike Vernal, Alana Goyal, Sonya Huang, Ramtin Naimi, Marc Bhargava, Cory Levy, Aashay Sanghvi, Konstantine Buhler, John Luttig, Varun Gupta, Ray Tonsing and Avichal Garg. Disclosure: I'm a small investor in American Terawatt, Convergent, Standard Intelligence and Trajectory (in this post), and in Factory, Physical Intelligence and SF Compute (elsewhere on the list). I didn't vote. The full list is on Breakout List.
6
3
106
8,311
Joan Cabezas retweeted
Sundial put the effort in getting RLVR right: training on verified fixes instead of artificial errors, smart rewards that disincentivize hacking, and deterministic scoring. The result — a trained Inkling-Small that fixes 83.7% of LaTeX errors in under 1 second and $.0013.
We fine-tuned @thinkymachines' Inkling-Small to fix LaTeX compile errors inside the editor, in under a second. On real errors it matches Claude Fable 5.1, six times faster and at about 1/40 of the cost. When we interviewed mathematicians, one of the most common annoyances we heard about in their day to day work is battling with LaTeX compilation errors. The most obvious solution is to ask a chatbot to fix errors. That works decently well, except switching to a different window breaks the writing flow, and copy-pasting the snippet often loses relevant context, especially when multiple files are involved. So we scraped 272,000 questions from TeX.StackExchange and kept the 3,978 whose snippet still fails today, whose accepted answer compiles, and whose difference replays exactly as a patch. The largest dataset previously released had 88. We tried tweaking the reward function in many ways. The first one was only a compilation check, which led the model to remove content until the document built. In the end, what worked best was rewarding compilation and a close match with the solution's PDF, and penalizing removed content. We're rolling out the model inside the @sundialmd editor this week. When a compile fails, the fix is applied as a suggestion in under a second and the PDF rebuilt. Write-up: sundial.md/blog/textinguishe… Thanks @tinkerapi and @thinkymachines for the support.
3
8
57
4,144
Joan Cabezas retweeted
We fine-tuned @thinkymachines' Inkling-Small to fix LaTeX compile errors inside the editor, in under a second. On real errors it matches Claude Fable 5.1, six times faster and at about 1/40 of the cost. When we interviewed mathematicians, one of the most common annoyances we heard about in their day to day work is battling with LaTeX compilation errors. The most obvious solution is to ask a chatbot to fix errors. That works decently well, except switching to a different window breaks the writing flow, and copy-pasting the snippet often loses relevant context, especially when multiple files are involved. So we scraped 272,000 questions from TeX.StackExchange and kept the 3,978 whose snippet still fails today, whose accepted answer compiles, and whose difference replays exactly as a patch. The largest dataset previously released had 88. We tried tweaking the reward function in many ways. The first one was only a compilation check, which led the model to remove content until the document built. In the end, what worked best was rewarding compilation and a close match with the solution's PDF, and penalizing removed content. We're rolling out the model inside the @sundialmd editor this week. When a compile fails, the fix is applied as a suggestion in under a second and the PDF rebuilt. Write-up: sundial.md/blog/textinguishe… Thanks @tinkerapi and @thinkymachines for the support.
13
10
126
11,434
Joan Cabezas retweeted
i love evals
29
38
736
25,045
Joan Cabezas retweeted
There is one proposal here that matters head and shoulders above the rest: Europe must have a stake in AI. This requires facilitating the construction of 20GW of datacenters by end of 2028, and 100 GW by 2030. Everything else is commentary.
Today, we publish a Transformative AI Strategy for Europe. Over the last few months, we’ve rallied researchers and engaged with governments to develop a plan for protecting the prosperity, sovereignty, and security of Europeans in a time of rapid AI progress. transformative-ai.eu @MonikaSchnitzer and I are honored to have convened an all-star team of thinkers and researchers contributing ambitious near-term objectives to make Europe relevant again, including: — Creating a Member State Alliance for Supply Chain Security (@anton_d_leicht et al) — Making European institutions ready to act in a transformative AI world (Conor McGlynn et al) — Securing Europe's share of global AI compute, in a ‘European Way’ that benefits local communities (@philip_fox_ et al) — Ensuring resilience to AI crises (@ben_s_bucknall et al) — Making Europe the global leader in assurance technology — And more objectives around security of supply and leverage (@milorignell et al), economic strength (@FraukeStehr et al), and safety/security (@NoemiDreksler et al) Our all-star senior expert council of Europe’s best and brightest (and non-European friends) reviewed drafts, provided strategic advice, and made suggestions for how to make the strategy more useful: @Ph_Aghion @Christophkw @bakkermichiel @ischinger @vestager @aleks_madry @FuestClemens Marta Kwiatkowska @LeoVaradkar @antonosika @DAcemogluMIT @Yoshua_Bengio It’s never been more clear: AI is real, and Europe needs to act. We’ve had all the warning shots and wake-up calls we need. Now the question is ‘what must be done?’ This strategy is our answer. transformative-ai.eu
5
7
66
7,636