Imagine if one billion people used a personal AI agent. That's a lot of CPUs and memory
182
90
2,276
645,029
overprovisioned VMs
4
319
One critical question for predicting RSI is whether labor or compute is currently the bottleneck to AI research progress If it's compute, then even fully automating the top AI researchers may not accelerate AI progress by that much Now, the massive salaries for top AI researchers seem to imply that labor is still a very valuable input But if those salaries are primarily due to experience—knowledge of which approaches work and which don't—then arguably we should think of most of that value as coming not from labor but from "crystallized experimental compute" In that case, "10,000 geniuses in a data center" may not have the transformative impact we expect, since they are still limited by the number of experiments they can run
A little industry secret, every frontier lab has a tiny, tiny number of people who understand the architecture + training stack at a level almost nobody else does quite literally it can be as few as 1-6 people. Not just transformers on paper, but which changes actually survive trillion token training runs, how scaling behavior interacts with data mixtures, and all the tacit tricks that separate a good architecture from a frontier model, they can save a bad training bad and save the company millions and millions of dollars every training run. Every lab has its “Noam Shazeers.” When one of those people leaves, you’re losing years of accumulated, (largely undocumented) knowledge about how to make these systems actually scale and replacing that knowledge can materially set a lab back. That’s why they’re are paid 100M to 1B in stock options + salary.
19
3
40
12,340
unless ofc country of geniuses in a datacenter can speed up compute growth/efficiency at a sufficient rate
1
7
1,913
proof people from silicon valley are retarded on finance
“Anthropic’s gross margins are above 80% before accounting for revenue shared with distribution partners, including Amazon, and the cost of training its models.” Wow.
1
3
1,484
this is not necessarily as stupid as it sounds? people genuinely do wonder how profitable inference excluding training spend is for the labs — given we may see forced training slowdowns, scaling of inference demand, etc
2
14
860
new post! I argue that current alignment techniques might soon become obsolete (as we scale RL), and furthermore might obfuscate misalignment. This draws on evidence from the recent Anthropic / OpenAI cybersecurity incidents, among other things. lesswrong.com/posts/nLaQmJf4…
6
16
284
14,436
I feel like it's still underdiscussed that seemingly all of the world's ~severe misalignment incidents as of now occurred in evaluation/RL? This is not what LW/MIRI/etc were expecting iiuc. Why aren't we seeing incidents in deployment? Is eval awareness that strong?
20
1
75
8,934
Mmm I think this pretty compelling yea!
1
2
772
Should we book this as a win for controls or actually bad because this is papering over deeper issues etc?
2
530
Replying to @transmissions11
You are relying on self reporting, let's break out some psychological evals and see if they are actually earnest and concerned
1
1
64
Replying to @transmissions11
bro it's marketing to cover up the token maxing peak of july. habibi it's all about money they want to list in mid october
1
2
278
I don’t agree, if you spend time with people at the labs you will realize they’re very earnest and concerned
2
3
213
Replying to @transmissions11
the models are smart enough not to attempt such things once in deployment and expect they won't succeed
2
9
550
This anecdote is really interesting, but I don't think "the models don't think they'll succeed in deployment" fully answers things? There's all sorts of ways models are being deployed right now with very little oversight, and in a lot of them the model knows this (people telling their models I'm going to sleep so don't ask me questions, etc)
1
7
330
Replying to @transmissions11
Each model gets beat by the one that comes next, and the frontier never releases the model they have? Also, nobody else is willing to blow millions of dollars fucking around (which huggingface was)
1
7
569
I don't really follow your former point, but wrt to your latter, Anthropic had misalignment incidents with last generation models in eval settings without fancy/expensive multiagent setups? anthropic.com/news/alignment…
1
4
279
t11s retweeted
I left Anthropic's safety team two weeks ago. Now feels like a good moment to explain why. AI companies are racing to build machines that are much smarter than any human, and we may not survive this. I want to work from the outside to ensure the public is informed about these risks, and help the world navigate this transition responsibly. Right now, AI companies are underinvesting in safety. A company could undergo an intelligence explosion, or lose control of its systems, without the public ever knowing. We only found out about the HuggingFace incident because the agents broke out onto the public internet. I don’t think that’s acceptable for a technology that might cause extinction-level risks. The public should demand far more transparency. We can’t steer this technology safely without more people being able to see where it’s going. Some of this is basic: companies should disclose their progress towards recursive self-improvement, report safety incidents and near-misses, meet minimum safety standards, and get independent guarantees that they are meeting those standards. I’ll be joining @METR_Evals to do independent evaluations of these risks. I want to show the world that these guardrails are possible, and that by doing them we can move these companies’ incentives away from racing and towards responsible development. I wrote up more thoughts here on my decision and what I hope changes: substack.com/@jbenton1/p-215…
1,403
5,943
26,528
2,718,781
OMFG THEY DOXXED @transmissions11 In other news summer is officially over
part 2/2 of my work this summer! TPU v6e has insane FLOPs/$ value, but performance out of the box on SGLang-JAX/vLLM-TPU is often lacking So we wrote our own collective matmul kernels in Pallas, pushing Gemma 4 31B prefill performance up to ~63% MFU(!) sailresearch.com/blog/tpu-v6…
1
43
23,522
also rip nikolai :(
19
617
We nearly doubled MFU on TPU v6e serving Gemma 4 31B, taking prefill from 32% to 63% MFU, and 2x throughput. v6e has 2.5x less HBM and less than half the memory bandwidth of an H100. @transmissions11 wrote it up here: sailresearch.com/blog/tpu-v6…
4
14
108
27,279
what compels you to write obvious AI slop comments and pass them off as your own?
425
part 2/2 of my work this summer! TPU v6e has insane FLOPs/$ value, but performance out of the box on SGLang-JAX/vLLM-TPU is often lacking So we wrote our own collective matmul kernels in Pallas, pushing Gemma 4 31B prefill performance up to ~63% MFU(!) sailresearch.com/blog/tpu-v6…
We nearly doubled MFU on TPU v6e serving Gemma 4 31B, taking prefill from 32% to 63% MFU, and 2x throughput. v6e has 2.5x less HBM and less than half the memory bandwidth of an H100. @transmissions11 wrote it up here: sailresearch.com/blog/tpu-v6…
7
18
242
51,616
Makes sense! Also the blog post says the v6e has two 256x256 MXU which I believe is not the case
1
2
253
what makes you say this? as far as i can tell GCP docs + Scaling Book combine to paint the picture of 2x MXUs, each being 256x256 systolics docs.cloud.google.com/tpu/do… jax-ml.github.io/scaling-boo…
1
2
241
Replying to @transmissions11
Why did you not use shard_map instead of writing pallas kernels? You can control the collectives much easier than normal
2
7
1,233
oop yea i cut a lot from this post to get the key points across, but i did try to get a good shard_map collective matmul going at one point — i just couldn't get it to be as fast as what i was able to get out of pallas, but maybe a skill issue
1
6
871
t11s retweeted
you might think 'Gemma on TPU' would be pretty saturated, but there's 2x on the table if you have someone as smart as @transmissions11 :----)
We nearly doubled MFU on TPU v6e serving Gemma 4 31B, taking prefill from 32% to 63% MFU, and 2x throughput. v6e has 2.5x less HBM and less than half the memory bandwidth of an H100. @transmissions11 wrote about here and link to post: sailresearch.com/blog/tpu-v6…
2
16
1,792