part-time researcher | scout @a16z | @ycombinator alum

San Francisco, CA
post AGI classics to read during the weekend: - Scenarios for the transition to AGI @akorinek - The Intelligence Curse by @luke_drago_ - AI Enabled Coups by @TomDavidsonX - The Ambition Singularity by @danfaggella - Gradual Disempowerment by @jankulveit - How long before Super-intelligence (1997) - The British Industrial Revolution in Global Perspective (2009) - Situational Awareness The Decade Ahead @leopoldasch - AI Monotheism vs AI Polytheism by @BerenMillidge
2
2
33
5,272
Joan Cabezas retweeted
Today we are open-sourcing 🥕Karotte, our framework for building RL environments. We've used it for the past year to build MLE RL environments for frontier labs, and it's been hardened through 1M+ evaluation runs and red-teaming.
30
44
304
26,507
Joan Cabezas retweeted
We’re on a mission to build a billion AI models, not trained by us, but by you. Introducing @bryel_labs we help your company build its in-house AI lab and give you the software to make model development and continual improvement simpler. We are backed by @ycombinator and have trained models at Apple, Framer, and F500s. We’ve enabled companies to train: • Classification models that are 20x faster • Search models that cost 5x less. • Long horizon agent models that match the frontier at 5-8x lower cost. Are you interested in partnering with us? Reach out. bryel.ai/
41
22
97
6,916
Joan Cabezas retweeted
We have seen a growing and concerning trend of startups supplying data to China. Today we are taking a public stance against this practice At fleet, we believe this represents one of the largest national security threats today and must be stopped. We must win the AI race for America and shape AI with the same democratic values that have shaped our nation Our full statement here: fleetai.com/supporting-ameri…
The AI arms race isn’t just being fought over Nvidia chips. Chinese labs are matching OpenAI and Anthropic by purchasing the exact same data from US vendors like @mercor and Surge AI (who also work with the US federal govt) forbes.com/sites/annatong/20…
4
19
121
10,111
Joan Cabezas retweeted
6-12 months ago the large labs started buying bio data this spend is starting to ramp very aggressively. the labs are somewhat indiscriminate buyers with enormous budgets. the therapeutics market has long had buyers of data (large and small pharma) but they've historically been the exact opposite of huge-budget and indiscriminate this shift will distort the bio startup market in a bunch of ways. there will be a lot of short-termist behavior to try to get in front of this capital firehouse. but unlike the AI data labelers (mercor, et al) and the robotics labelers (mecka, et al), the bio startups serving up this data have a chance to emerge from this period with businesses that don't rely on selling data the best bio startups will use these sudden resources the labs are dumping on them to supercharge their novel assay and work towards independently powerful models and eventually directly produce therapies bio is the final frontier and its a very exciting time ahead for bio startups IMO
42
46
898
132,401
Joan Cabezas retweeted
We investigated a seemingly simple question: how do you grade if AI generated parts or assemblies are correct? Here's how we improved on this with our verifier, FrontierCAD.
9
14
328
94,197
Joan Cabezas retweeted
Backprop has been the only credit assignment algorithm capable of training large neural nets. Introducing Dust, the first zeroth-order method to pretrain transformers to approach and sometimes even *exceed* backprop with large amounts of computation (bitter lesson!) - Dust is on the order of 1,000 to 10,000x more compute efficient than EGGROLL, the state-of-the-art ES method, for training transformers. - We chose pretraining transformers because it's arguably the hardest task possible for zeroth-order optimization. - Search in high dimensions is really misunderstood: zeroth-order usually gets better with overparameterization, not worse, i.e. at a fixed population size, larger models (up to 120x) reach a lower loss than smaller ones. - Backprop-like gradients emerge from a large population of activation perturbations, and the alignment holds up at every scale we tested, up to 1B tokens, which is critical for scaling. The core idea is a *virtual population*, which helps scale to large populations in transformers by perturbing activations in parallel. w/ @bishmdl76, @cs_serdar, @akshayvegesna
59
278
2,476
231,502
September was Terac’s best month yet. vs. August: Expert earnings: +248% Experts earning $1K+/month: +196% Experts referring others to Terac: +84% All from an extremely small, talent-dense team. Every improvement means more people matched with better opportunities, more experts paid for what they know, and more companies finding the talent they need. Our mission is to efficiently allocate the world’s labor market. We’re urgently hiring exceptional operators and engineers to help us build it.
3
17
870
Impressive team, you should consider joining them!
1
99
Thought Jev is fast? We evaluated Jev on cuaspeedrun.com for computer use agents and found: 1. 1.5x slower than Opus-5.5-xHigh 2. 10x lower accuracy 3. 5x more steps That puts it in the GPT-4 league of performance and Astra-xHigh on time What and Why?? 🤯 🧵
5
9
53
12,168
Joan Cabezas retweeted
Building AI inside your company? These are your people. Two @bryel_labs events at SF @Techweek_ next week for founders, CTOs, and engineering leaders: Tuesday evening: Join me and @vinitra_s from @ScholeAI to discuss what it takes to level up your intelligence within your company. (Co-hosted by @thehousefund) 🗓️: partiful.com/e/wmdrrEPxrjo20… Thursday afternoon: Curious about training your own models? Building your intelligence stack? Come swap notes with other pioneers making the same decisions. 🗓️: partiful.com/e/BgX5jT0yxM3fa…
2
5
19
264,307
Joan Cabezas retweeted
Today, I am excited to announce Animation Bench. Our first benchmark to evaluate frontier models on web animation reconstruction. Frontend generation is a common commercial use case for coding agents. In addition to static layout it contains transitions and interactions triggered by dynamic behaviours. We evaluated four frontier models on 48 tasks from real websites across three axes: 1. Visual Similarity 2. Motion Consistency 3. Layout Correctness. We find that motion consistency is often incomplete or missing.
Introducing Animation Bench, our first benchmark to evaluate frontier models on web animation reconstruction. Coding agents can already recreate visually plausible web animations, this work evaluates if they can reconstruct behavior of a web page.
22
16
255
20,500
Joan Cabezas retweeted
Introducing Synthetic Hospital: an open, fully synthetic longitudinal EHR benchmark with verifiable ground truth! 1,268 patients, 5,602 encounters, zero PHI. Physicians could not reliably distinguish its charts from real ones. 📄 arxiv.org/abs/2609.30027 💻 github.com/sparkcpark/synthe… ✍️ sparkcpark.github.io/posts/f…
54
146
1,306
299,178
Joan Cabezas retweeted
In 12 weeks, we built a research facility that is run entirely by AI. AI designs, executes, and observes experiments end-to-end across biology, chemistry, and materials science. We’re introducing SciUniverse: a benchmark that measures AI’s ability to do real-world scientific research.
241
397
2,795
790,786
Joan Cabezas retweeted
There’s a good chance your open source model is costing more than the frontier. Cheap tokens ≠ cheap tasks. Here, we introduce Intelligence Density, and Density Aware Training, our post-training technique to achieve less wasted compute, better learning, all with no knobs to tune. Enabled by default in every Trajectory model.
23
24
287
56,475
I see sooooo many companies living off the RL data firehose. Labs, due to their distaste for getting their hands dirty with data, are creating a whole outsourcing industry. The results won't be optimal. Post-training data is not a commodity, unlike its complement. A lot of thought has to go into building environments worthy of gradient descending. Labs not taking more ownership of the data supply chain makes me more bullish on vertical AI.
4
3
114
7,423
aren’t frontier labs doing this in house now?
2
1
584
Joan Cabezas retweeted
More people need to know about this lab. qlabs.sh/
5
4
106
15,758
The poaching offers that non-technical human data ops leads (SPLs) are getting would blow the minds of 99% of people. Many of these are at parity with MTS offers at the labs. How long until the music stops is a different question...
People think the data wars are winning a big deal with a lab, when really it is everyone trying to hire from one another
9
2
116
52,702
Joan Cabezas retweeted
My new report with @luke__emberson.
AI is getting cheaper more quickly than any other transformative tech in history. At a given level of performance, cost has fallen ~47%/quarter since 2023. That’s 4× faster than DNA sequencing, 6× faster than compute, 18× faster than lithium batteries, and (up to 1973) 54× faster than electricity.
2
8
53
3,414
Joan Cabezas retweeted
Reinforcement learning is having a moment. For the past few months, we’ve been working with @tryterac to put together something special: Convergence — an event for the people building the future of RL. Speakers from @GoogleDeepMind, @databricks, @mercor, @arena, @joinHandshake, @turingcom, @deeptuneai, @contra, @Stanford, @Princeton, @UCBerkeley, and more. Our goal: make this THE event for the RL community. Oct 29 in SF. Come join us → terac.com/events/convergence…
13
8
77
21,241
Joan Cabezas retweeted
Trajectory is at @HackTheNorth, come say hi!! We’ve got some cool merch in store for whoever can come up with the wackiest benchmark ideas
6
3
64
3,187
Joan Cabezas retweeted
Introducing WeirdML v3, a fully agentic benchmark featuring 11 complex hand-made tasks. Models must explore and understand unfamiliar data, develop ML and data analysis pipelines and produce results despite limited data, unspecified goals and/or very limited feedback. 1/8
WeirdML v2 is now out! The update includes a bunch of new tasks (now 19 tasks total, up from 6), and results from all the latest models. We now also track api costs and other metadata which give more insight into the different models. The new results are shown in these two figures. The first one shows an overview of the overall results as well as the results on individual tasks, in addition to various metadata. The second figure shows cost vs performance and shows a clear scaling with better results for higher costs. We also have a very varied pareto frontier with 11 models from 6 different companies having the best accuracy for a given cost for at least some of the cost range. Grok 3, Claude Opus 4 and GPT 4.5 are the ones that underperform for their costs, while Gemini pro and o3 pro have the best results at the highest costs. Qwen3 30B3A, grok 3 mini and deepseek R1 also each represent a good chunk of the pareto frontier.
54
56
655
140,668