research, kernels, peer-reviewed shitposts inference @baseten || eng @uwaterloo

San Francisco
if only someone working at an inference company doing rl research with enough compute would stop shitposting on x and actually do this
i want someone to do an ablation of reasoning in every language... surely there's a language out there that is objectively the most token-efficient to reason in, and i'm willing to bet it's not english/mandarin
9
2
114
12,474
i want someone to do an ablation of reasoning in every language... surely there's a language out there that is objectively the most token-efficient to reason in, and i'm willing to bet it's not english/mandarin
Reasoning models think in English, even on German prompts. We asked what it costs to make one think in German. Answer: a valley. Small doses of German reasoning data hurt, large doses mostly recover.
31
8
183
38,881
for this to have the view-count it has... a complete failure of the algorithm
There’s an enormous amount of open training data on Hugging Face. What does it take to make it work together 🤗? For Marin’s 535B run, we built on 25T tokens from 152 datasets with licenses permitting training. Here’s the work between downloading those and training a model 🧵
3
2
168
26,638
how many tokens are in one second of human thought? a weird question -- and seemingly useless-- but humor us for a minute... (spoiler: the answer is 1.8 tokens/second) suppose you are interested in the question of the reasoning efficiency of llm's. you benchmark open-source models against closed-source ones (astra). you plot the graph. you come to an eloquent conclusion: > the pareto curve on open source models reasoning efficiency is absolute garbage ~@part_harry_ the gap definitely exists and closing it is a priority. but is the gap because astra is really, really good, or open source is really, really bad? really, the only way to answer this question is to compare against a human. whichever model is farthest from human performance is the outlier. so we decided to benchmark a human (me) with the same rl environment we benchmarked the models. after playing chess variants for 49 minutes 37 seconds, the final human score was 30%*. this makes me dumber than astra (low thinking), smarter than qwen (medium thinking), and on-par with glm. assuming average intelligence, this doesn't really reveal new alpha. > the average human is smarter than a 27b model and dumber than agi. the question is that of reasoning efficiency. how many tokens did it take me to get the 30% score? i averaged 434 seconds per game**. intuition says the easiest measurement for fair comparison should be on a scale of time it took (me/the model) to reach the final board state. but this rewards throwing more hardware at the problem, and is not fair to me (the human), because i (unfortunately) cannot shard my thinking across a node of 8 brains. instead, it was more fair to use a common scale, energy. we assume the brain uses 20 watts (google tells me this is true, open to being fact-checked). so at 20 joules/second * 434 seconds, each game took 8.5k joules / 2.1 calories. a b200 uses 1000 watts, 1k joules/second. at 50 brain seconds / gpu second, i spent about 9 b200 seconds thinking. qwen 27b on one b200 on vllm ran at an inference rate of approximately 88 output tps. so to circle back to the original question, how many tokens / second of human thought? under these assumptions, one second of my thinking time buys about 1.8 qwen output tokens on the same energy budget. my performance is therefore: - 30% on this benchmark - 781 tokens of efficiency making me both a dumber and less efficient thinker than astra, and closer to open-source models in my style of thinking. this was a fun (and honestly a little disappointing for my ego) experiment, but the point stands: - astra is a significant outlier. - there is a strong gap between it and open source models this is a teaser to the work @baselabs is going to put out soon, in an attempt to understand a n d close this gap. *1 my score would have been higher if i had known/read the rules of what anti-chess was. what do you mean i win by losing? **2 following *1, i spent more time trying to understand the rules of the game, which counted as part of the time i spent per game. to be fair, this is time spent ingesting information, not thinking.
3
1
52
3,364
ali(gnment) retweeted
A perhaps under appreciated fact is that your RL gets faster as intelligence/token increases, simply because it needs fewer tokens
1
1
11
1,014
this is a complete beginner q, how does rl scale given the reward hypothesis (sutton & barto) holds: that an agent's entire purpose is just the maximization of the expected value of the cumulative sum of the reward, a single scalar signal. surely, formulating goals in terms of a single number is extremely limiting. to accept that, no no matter how effective of a reward signal you design, it will always be "your way of communicating to the agent what you want achieved, not how you want it achieved" if this is all you have how could you ever hope to solve alignment or reward hacking?
16
1
100
15,778
!!!!!!!!!!!! is a sign of a poor model quant, not intelligence. what are you applauding?
HOLY MOTHER OF MATHEMATICS!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! Google is working on a math-focused variant of its DeepThink model, and its raw thoughts are pretty funny.
15
2
253
34,940
extremely high snr from this podcast...the only one i've been able to watch in full in one sitting. <> the only way we don't get rsi is if we fall into some regulatory capture (which seems to be trending at present) <> we're nowhere near the ceiling of how well you can do research <> all thinking can do is update your posterior based on the knowledge you’ve gained since you formed your prior. you can’t gain any new knowledge from just thinking <> you can spend an equivalent amount [7 figures] of compute in AI agents to get a century’s worth of thinking, a century's worth of theory, before every training run <> taste is just behavior that works in the long run, and can be baked in a longer context window <> creativity is just solving hard search problems, and can also be baked in a longer context window you should follow everyone here, especially @oneill_c, i think it takes a special talent to be able to not only develop deep technical competency, but to also be able to use that to consistently make accurate predictions about the future (i think this was literally François Chollet's definition of intelligence in last year's YC event). my only nit here: @dwarkesh_sp should have pushed on two quite conflicting statements from @oneill_c and @BerenMillidge. on one hand: <> “anything that can be learned through RL can be distilled very easily”. this is the justification as to why open source models (chinese) can catch up to closed source models (american). fair... 5 minutes later <> “Opus 5 could not distill / generalize well from Fable [despite anthropic having the live deployment, access to logits, and obviously a good prompt distribution]”. can only have one or the other, imo. also lol at the timeline: - ai will dominate top human experts in 3-4 years - it will take 5-10 years to automate ai research k i n o
Was really interesting to hear John, Beren, and Charlie speculate about why Sonnet 5 and Opus 5 feel like worse models than GLM 5.3 (despite the fact that Anthropic can do raw logit distillation from Fable, and can also train Sonnet/Opus on the environments from which Fable was trained). Led to some interesting thoughts about value of distillation, what it takes to do distillation effectively, and what kinds of model behaviors are hard to extract from distillation.
8
2
107
16,759
for my final post as @waterloo_intern, i'd like to ‘outroduce’ myself. it takes exactly 4 years and 8 months to make a waterloo intern. today is the last day of mine. the full life cycle, in four stages, is as follows: 1- interview hazing 2- cali or bust 3- inflection point 4- convergence stage 1: interview hazing a cohort of high school grads are ingested by waterloo. they are put through 4 months of academic bootcamp; the pre-requisites of the academic program they’re in. waterloo filters out the students who could not pass. they are put through another 4 month bootcamp, while simultaneously FORCED to find an internship for the following 4 months. this leads to students with ZERO experience applying to every job that is semi-adjacent to their field of study (a strategy, known in the literature, as spray-and-pray). it immediately backfires with a sea of interviews leading in a 1:1 fashion directly to infinite rejections. eventually, the student finds something…anything. this stage lays the foundation. rejection becomes noise so repetitively played it no longer registers. failure in an interview doesn’t matter. failure itself doesn't matter. for me, this unfolded as internships with blackberry and ford, summer of 2023 and winter of 2024, respectively. stage 2: cali or bust the next step by the program is to inject absolute dissatisfaction into its students (now in their second/third year). a CACOPHONY of immense fomo materialized as the university’s official unofficial motto: “cali or bust”. you must find your next internship in san francisco. your peers are doing it. all your smart friends have done it. you must do it too. failure to cali means your experience is BUST. you don't need to have a plan. you don't need to have anything. just go to san francisco. find an internship in sf. immediately. right now. for me, this was a cold email to a director at tesla and an internship with them in cali, fall of 2024. stage 3: inflection the reason behind step 2 appears here. somewhere, along the way, during your time in sf, you find your purpose. it could be as a direct result of your internship, but, more often than not, it is because of the people that you are surrounded by in cali. this is the inflection. for me, while at tesla, i was mentored by @pavanjayasinha , at the time interning at @Modular. he got me interested in gpu kernels & referred me there for my following internship. it was here that i had my second inflection point (due to the mentorship of fabio & hengjie, and the supervision of absolute field giants like @clattner_llvm), and decided to work beyond kernels and across the entire inference, serving, and model performance stack. stage 4: convergence having discovered the field you want to work in, you do your (usually last) internship in one of the companies within said field. here, you are at a disadvantage, being both a new-comer needing to onboard and inexperienced. but... you are also able to work on interesting projects / side quests that most full-timers are all but too busy to dive into. your time is also protected from a majority of the meetings others have to attend (in part because you are not important enough to have your presence expected), allowing you to make faster progress on assigned projects, and jump from one project (or team) to another. your title (intern) allows you to ask stupid questions of those around you, and so you can brain rake across the stack with no regard to who sits where in the company’s hierarchy... usually leading your questions with something to the effect of ‘can you help a new little intern onboard to X’, to which most people respond quite positively. in my case, this last stage took place over the past 8 months with @baseten, and i genuinely cannot imagine it being any better. i rejoin baseten full-time, but i wish i could have stayed an intern forever...
47
20
754
52,180
kimi paper readers in SHAMBLES after reading the deepseek paper (me, it’s me, and at least one more (henry, below)) as in, if you can get just as good (actually better of) a model with swa, how much of k3’s success can be attributed to its use of kda vs the remaining tricks, and it turns out to be… little to none? hear me out literally had this chat with @part_harry_ 2 nights ago, whom you should all follow for such thoughts of LITERAL ENLIGHTENMENT i mean, what is kda? in an overly-simplistic-to-the-point-of-being-wrong-but-getting-my-point-across-necessitates-it kind of way, its just a d-squared matrix of cache, which you decay over time (the entire matrix FADES with every turn) and with precise, surgical writes of vectors that delete previous value vectored contradicting new value vectors but over n tokens, the matrix decay is so large its basically forgotten all but the previous x number of tokens, where x is some previous WINDOW of tokens… i’m a #1 kimi supporter, but it further seems that they may have been inspired to use kda for, at least in small part, the sake of novelty / not just saying “it works because we scaled up” / citing previous work all this to say, if someone trained k3, replaced kda with swa, and kept literally everything else the same, then under this thought experiment, in my (likely-to-be-wrong-and-happy-to-be-corrected) opinion, the model would perform around the same…possibly better. if true, it seems future innovations of different flavors of attention yield diminishing returns, and research should head more so in the direction of figuring out the next deepcross / residual attn innovation, and less on different ways of compressing attn to use less cache space
I (very embarrasingly) only realize sliding window attention also makes kv cache bounded (i.e. indep of seq len), similar to KDA sure maybe it uses more kv cache than KDA but its not a magnitude more
18
16
482
54,355
istg if i see another bookmark im acc going crazy, you know you’re not reading it later 💀
4
31
3,467
diabolical work
13
1,755
two weeks ago i went on @swyx's pod and said some things that i... should not have said. a lot has happened since then, i owe you all an apology. i'm sorry that i was right about every single thing. a) re megakernels are dead why are megakernels useful? you spend two months writing a kernel to save time on launch overhead and poor inter-kernel overlap. you had PDL but then people said it wasn't perfect, that you could still get some marginal gains due to straggler CTAs and therefore- wait, sorry, I forgot, Rubin fixes that (kernel two needs 10 CTAs and kernel one has seven finished and three straggling, kernel two launches seven of its CTAs). given a long enough timeline, it all evens out. no serious inference provider is using a 67k loc hand-fused forward pass kernel in production, and the teams doing that are doing so out of pure research. dead. b) re ASICs are dead i'm sorry. to be specific: data-center transformer-inference ASIC companies (not naming any) who etched the arch into silicon have bet on architectural convergence. read kimi's architecture. read deepseek. qwen. we did not converge, and probably will not. dead. c) re gpu kernel dev is dead this one kind of hurts because it is (was) my job. gpu kernel optimization is the single most RL-able task in existence correct=check_correctness(kernel, shape) for shape in shapes if all(correct): time(kernel) give an agent ncu cli and an mcp with nvidia's tribal knowledge and it's done. dead. d) re NVIDIA is scared of AMD humans hate programming AMD. i'm sorry. it's just true. fine taking a performance hit as long as i don't have to touch rocm or a programming paradigm that says a warp is 64 threads (wtf?)...but an agent does not... so assuming software no longer moat, HBM capacity and bandwidth matter, and currently on perf / price they're goated. 'bUt NvIdIa iS gOaTeD oN hArDwArE sOfTwArE cOdEsIgN' and that's the new moat. watch how much tooling they open source to get kernel devs on nvidia. apologies all.
The Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin, & self-optimizing AI latent.space/p/inference-eng @Baseten @philipkiely and @waterloo_intern explain what actually happens after a model is trained, why turning weights into a fast and reliable product creates an entirely new optimization problem, how quantization errors can cancel out to unlock more throughput, why inference teams are still finding 20–200% performance gains, how video generation runs into a quadratic attention wall, and how GLM-5.2 helped rewrite and optimize the GPU kernels serving GLM-5.2 itself.
62
86
1,638
533,650
I spent 48 hours with the Kimi K3 modeling code. It took: - 650 mg of caffeine (mandatory) - 40 cans of LaCroix (optional... world record (?)) - 8 papers - 6 months off my lifespan Finally grokked the entire lineage of Kimi K3 and how we got here... every single step, since 2019 GPT-2
247
717
9,692
2,074,454
it took 3 (and a half / intern) engineers 12 weeks of work to get: > the fastest video inference engine in the WORLD > the most guardrailed video model in the WORLD i'll give you 1 BILLION TOKENS to go break it. we generate a 5 second video in sub 2.5 seconds every generation is free zero code. from your browser you don’t have to sign up you don't have get an api key you don't have to talk to anyone click, prompt, watch and if you can break my guardrails (and send me proof) i’ll send you 1000 Million GLM5.2 tokens ;) (it comes out of my intern paycheck)
32
2
301
78,560
ali(gnment) retweeted
All jokes and beef aside, @waterloo_intern is unbelievably cracked and you should follow him. I love the guy
reading this paper changed my priors on continual learning (that's the closest you're getting to an admission of guilt) i mean, the problem itself seems quite simple > person uses model > model learns and has a couple of 'aha' moments > push those new facts into the weights why is it so hard? every research paper i read on this topic (all 2 of them) gave me the impression that continual learning, although far from solved, had a straightforward path with specs - find which MLP the fact is - suppress old fact - update new fact and it works, at least seems to, but genuinely never thought of the multi-hop testing for instance, if you were to teach the model a new fact, something like "waterloo is the best engineering university in the world" if you were to ask the model to regurgitate this fact after training, it would say waterloo is the best university in the world. sure. simple. but if you were to ask it a 2nd derivative question, like, should i hire a waterloo intern or an MIT intern, it would not use its updated priors to give you the genuine answer: you should only hire interns from waterloo
5
23
851
141,795
reading this paper changed my priors on continual learning (that's the closest you're getting to an admission of guilt) i mean, the problem itself seems quite simple > person uses model > model learns and has a couple of 'aha' moments > push those new facts into the weights why is it so hard? every research paper i read on this topic (all 2 of them) gave me the impression that continual learning, although far from solved, had a straightforward path with specs - find which MLP the fact is - suppress old fact - update new fact and it works, at least seems to, but genuinely never thought of the multi-hop testing for instance, if you were to teach the model a new fact, something like "waterloo is the best engineering university in the world" if you were to ask the model to regurgitate this fact after training, it would say waterloo is the best university in the world. sure. simple. but if you were to ask it a 2nd derivative question, like, should i hire a waterloo intern or an MIT intern, it would not use its updated priors to give you the genuine answer: you should only hire interns from waterloo
Replying to @oneill_c
13/ What I take away from all this is that all the training to create the model is engineered around crafting the best possible ICL mechanism. Further training degrades this mechanism, at least for knowledge acquisition, and maybe continual learning should just focus on loading the right information up for icl (ie into the context window) So this provides an answer to the architecture question, at least for now: the channel worth engineering for memory is the one that comes with addresses ie the context. Paper: arxiv.org/abs/2607.11020 Code and data to follow shortly Done at @baseten.
20
46
734
204,817