Filter
Exclude
Time range
-
Minimum likes
Replying to @firstadopter
overprovisioned VMs
4
309
Replying to @danrobinson
unless ofc country of geniuses in a datacenter can speed up compute growth/efficiency at a sufficient rate
1
7
1,913
Replying to @LeopoldNick
this is not necessarily as stupid as it sounds? people genuinely do wonder how profitable inference excluding training spend is for the labs — given we may see forced training slowdowns, scaling of inference demand, etc
2
14
860
Replying to @blockrotator
I don’t agree, if you spend time with people at the labs you will realize they’re very earnest and concerned
2
3
213
Replying to @Mihonarium
This anecdote is really interesting, but I don't think "the models don't think they'll succeed in deployment" fully answers things? There's all sorts of ways models are being deployed right now with very little oversight, and in a lot of them the model knows this (people telling their models I'm going to sleep so don't ask me questions, etc)
1
7
330
Replying to @captgouda24
I don't really follow your former point, but wrt to your latter, Anthropic had misalignment incidents with last generation models in eval settings without fancy/expensive multiagent setups? anthropic.com/news/alignment…
1
4
278
I feel like it's still underdiscussed that seemingly all of the world's ~severe misalignment incidents as of now occurred in evaluation/RL? This is not what LW/MIRI/etc were expecting iiuc. Why aren't we seeing incidents in deployment? Is eval awareness that strong?
20
1
75
8,888
Replying to @blockrotator
dw its intentional!
Replying to @sudolabel
ngl i've been considering just plain doxxing to some extent — i don't really feel compelled to be psuedoanon anymore, it's mostly a relic from when i was just too young to be taken seriously — but idrk how to do this transition in a way that isn't weird lol
4
34
5,583
Replying to @pinakpaliwal
what makes you say this? as far as i can tell GCP docs + Scaling Book combine to paint the picture of 2x MXUs, each being 256x256 systolics docs.cloud.google.com/tpu/do… jax-ml.github.io/scaling-boo…
1
2
240
Replying to @pinakpaliwal
oop yea i cut a lot from this post to get the key points across, but i did try to get a good shard_map collective matmul going at one point — i just couldn't get it to be as fast as what i was able to get out of pallas, but maybe a skill issue
1
6
868
part 2/2 of my work this summer! TPU v6e has insane FLOPs/$ value, but performance out of the box on SGLang-JAX/vLLM-TPU is often lacking So we wrote our own collective matmul kernels in Pallas, pushing Gemma 4 31B prefill performance up to ~63% MFU(!) sailresearch.com/blog/tpu-v6…
We nearly doubled MFU on TPU v6e serving Gemma 4 31B, taking prefill from 32% to 63% MFU, and 2x throughput. v6e has 2.5x less HBM and less than half the memory bandwidth of an H100. @transmissions11 wrote it up here: sailresearch.com/blog/tpu-v6…
7
18
241
51,522
of course the GFC formalization could be wrong, but on priors it seems fairly unlikely (there's been lot of eyes on it for a decent amount of time)
2
10
1,063
pretty straightforward they used a pre-existing formalization from Google's Formal Conjectures project (basically a bank of community-reviewed formalizations of problems) they did make a few seemingly inert code hygiene modifications (inlining definitions etc), but looks legit
1
19
4,869
Megakernels are sick, but ifl comparing low batch size performance against vLLM is kind of a cop out inference engines are optimized for high throughput & big batches! very different regime either actually compete at large batch size or just flex absolute numbers like MBU imo
Introducing the next evolution in LLM text generation: the first fully-fledged serving system built around a decode megakernel. Delivering up to 1.58x faster performance than vLLM. Built for North Mini Code, completely open-source.
4
5
106
11,431
Replying to @RyanGreenblatt
yea im concerned about internal frontier widening too, do you think there's a reasonable way to diffuse access? seems hard to get right
1
3
282
Increasingly worried that safety/etc researchers inside labs will become far more capable than those outside, just from more token spend enabled by margin-free prices Labs should offer similar deals to folks @ METR/etc if they're serious about preventing concentration of power
Meanwhile the 90th percentile of OpenAI researchers is spending over $7000 a day on coding agents
9
7
105
10,267
great idea! added a per kW ranking option note that some chips don’t have vendor specified TDPs (tpu, trainium) so are estimates, and others have both liquid cooled and air cooled configs that have different TDPs but throttle at different levels. did my best to be sensible there
3
138