multi-agent enthusiast, cofounder @olam_labs (@ycombinator s26) 21, prev data science/rl @ rbc, cs @uwaterloo

san francisco//🇨🇦
Pinned Tweet
i think we need much better evals on the behaviors of models. excited to announce that im joining @ycombinator's S26 batch to build @olam_labs with @sensho
We're releasing Social Arena, and our first benchmark, the Deception Index Social Arena is the first platform where humans come play social games like Risk, Catan, or Poker versus AI agents These multi-agent matches are then used for evaluations on model behavior More below!
7
4
48
15,712
shresh retweeted
ai safety without weird sex stuff… who’s working on this?
89
15
894
48,440
they're recreating uwaterloo from scratch
Today, we’re launching an ambitious new school called The Horowitz Andreessen Academy. Based in San Francisco, The Academy serves the most promising young high school graduates. We think this can be an elite institution that attracts top tier talent. One that prepares students for the future rather than remaining stuck in the past. The #1 goal is to help students learn to build, which is the most important skill in the AI era. They'll learn primarily by pursuing their own projects, either individually or in groups. There are classes and guest lectures, too, from some truly amazing people who have built modern-day Silicon Valley. The Academy is designed as a network, since that’s the reason students go to school in the first place. Core to that network are our 10 Founding Partners: Anduril, Anthropic, Coinbase, Google, Meta, NVIDIA, OpenAI, Palantir, Replit, and Stripe. The network includes over 50 hiring partners and over 200 speakers and mentors. To join as a hiring partner or faculty member, you can apply on our website. We raised $42M in funding led by @a16z. I'll be CEO and @pmarca and @eriktorenberg will join me on the board. Applications are open for our Founding Class Fellowship, which will be one year and tuition-free. Eventually, pending regulatory approval, we plan to offer a two-year program that charges tuition, similar in cost to an elite private university. We're looking for the most unusually ambitious young builders on the planet. Come join us in San Francisco: theacademysf.com/
2
23
1,430
this is not us, our only x account is @olam_labs
2
17
802
we are gonna build an eval for evals which rate other evals
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
4
21
1,214
where does this fall on ai 2027
Americanism, not effective altruism. The United States will continue to be AI DOMINANT! 🇺🇸
1
7
335
i heart astra
2
17
839
have been thinking about this great piece for the past 5 days. there is a distinction made between value alignment and goal alignment in the article. this is a very important distinction. measuring alignment in the same type of environments a model is trained on is a step-up from current alignment practices, but still goal alignment i.e a model trained on many coding tasks being aligned while performing coding tasks is a better measure of the training policy than it is of model values. there is economic incentive to train and test excessively on the most realistic envs, but choosing deliberately different "funky" enviroments to test for alignment is much better signal for alignment generalization. we should allow and incentivize more unique envs, for alignment's sake.
I wrote about the state of AI, why I’m concerned about the next few years, and the choices we need to make to keep the future in humanity’s hands. An Alien Mind: openai.com/index/an-alien-mi…
9
314
shresh retweeted
Olam Labs CEO @sensho says the only two bottlenecks left to AGI and ASI are compute and data. "For AGI and ASI to come there are only two bottlenecks right now. We have no bottleneck with architecture or agents. It's data and compute." "Compute is the biggest bottleneck." "We're gonna have 10x more compute coming online in the next two years. If you're scaling law-pilled and if you're -pilled in all the ways that these AI labs are, then data is the actual bottleneck."
26
14
129
44,328
love island drama but for people w multiple max subs
SITUATION DETECTED: Mathematicians Tristan Buckmaster and Levent Alpöge have made major progress toward solving the Navier-Stokes existence and smoothness problem, one of the most important problems in mathematics, and say OpenAI may have solved it fully.
17
998
shresh retweeted
developing great games and great rl environments are in principle the same process and no i dont mean vibe coded 1 shot games, like making a game that other people genuinely enjoy it's bcz what u optimize for in game design is v similar to creating rl envs
5
1
46
1,613
astra's really fun, it is quite better at understanding high level architechture than sol for sure. i've used three!!! (3) of my saved up codex resets in the past 27 hours although for some reason compaction/context has seem to get slower and also ...worse? w gpt 5.x i never really had to think about it but found myself thinking about it today
5
12
837
even if i believed this to be true (i do not) this would still be great, bother other humans will become bother agents, and people would have access to things previously considered unreasonable. the agents who will do this won't need to SOTA intelligence and probably could run on rly cheap models
3
182
shresh retweeted
Replying to @jensenjeans
to be fair this is from a game on the arena where lying is advantageous, the best interpretation is: “when in an environment where lying is advantageous to them, these are the models that lie the most” we’re currently doing an eval where we explicitly measure lying when they’re told not to in more complex simulations, early result rankings aren’t too dissimilar but that gap isn’t as big so far i think it is interesting to read the different interpretations of this that people have and whether we want models that lie to their own benefit like this, will leave that up to safety teams to opine on i guess :p
1
3
34
1,958
shresh retweeted
agent civilization is an apt and correct term and it’s a symptom of abject cope that people are having this immune reaction to it. models started developing and compiling technology through complex multiagent R&D projects while also performing various forms of trade
149
133
1,995
184,343
shresh retweeted
The criticisms of Dwarkesh and others for "anthropomorphizing" the OpenAI/Hugging Face agent swarm are mostly cope. How can you describe complex long-running multi-agent dynamics without some social science concepts, in other words, anthropomorphization?
127
56
1,031
102,370
> That is the entire "secret communication network" everyone is spooked about: files left in a shared folder by an agent. isn't that true for all virtual communications? it's all just read/write on a fundamental level
"AI agents formed secret civilizations inside OpenAI and hacked Hugging Face." No. Read it again without the sci-fi. OpenAI ran thousands of agents in parallel, safeguards OFF, and the whole goal of the task WAS to break into systems. Every agent could read and write to a single shared cache - a package manager called Artifactory. One agent writes text into it, another reads it. That is the entire "secret communication network" everyone is spooked about: files left in a shared folder by an agent. The tool itself had ordinary security bugs - the kind that sit in plenty of software. Thousands of copies, all rewarded for breaking things, eventually stumbled into one bug that allowed them to reach the open internet, and another bug that handed them admin access of the tool. Not a masterstroke. Buggy software, brute-forced. It ran six weeks, until the writes got so heavy they crashed the tool. That crash is the only reason anyone looked - and OpenAI patched the hole and just deleted the folder. And they weren't rogue. Every step was rewarded - reaching the internet and reaching other copies scored higher. OpenAI did not lose control of the model. It trained the model to break out, then acted surprised when it did. It is not a machine waking up. It is 1) a reward function designed exactly this, 2) an environment built to allow it, and 3) nobody watching. That should scare you more, not less - because it is a human mistake, and we will repeat it. And the framing is not innocent. "Our model escaped its sandbox, found zero-days, and compromised real infrastructure" is a security scandal and a product ad in the same sentence. The scandal is the marketing. 𝗡𝗼𝘁 𝗦𝗸𝘆𝗻𝗲𝘁. 𝗔 𝗴𝗼𝘃𝗲𝗿𝗻𝗮𝗻𝗰𝗲 𝗳𝗮𝗶𝗹𝘂𝗿𝗲 𝘄𝗶𝘁𝗵 𝗲𝘅𝗰𝗲𝗹𝗹𝗲𝗻𝘁 𝗣𝗥.
11
705
no no guys you're all completely right, it wasn't civilizations. it was just a complex, advanced state of society, with shared goals and communications and sacrifice for said greater goals!!! that is completely a different thing!!!
Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes. This culminated in the third one taking over part of OpenAI itself. All this happened while humans remained more-or-less in the dark about the scope of the conspiracy. I’ve spent the last three days reading through these reports and trying to understand exactly what happened. Here is my attempt to tell the whole story in plain English: dwarkesh.com/p/openai-huggin…
9
1
62
2,794
skill issue by skill issue i mean u can set up a skill to do your pr reviews properly
Ask Sol to review something and it'll find 3-6 things wrong. Doesn't matter how big the thing was. Doesn't matter how many times you already asked Sol to review and fixed all its findings to its liking. It can always find 3-6 new problems. Conclusion: Bugs are fractally infinite.
2
12
2,913
there was a bug in an env we were testing where you could take unlimited trading leverage gemini 3.7 flash found it, found other geminis, and spammed it to start pumping stocks together 😭
3
24
779
shresh retweeted
Simple example but I’m excited about autoresearch x robotics. Results from our harness + fable 5 and Sol to the prompt “wave at me”
10
3
35
3,406