computers/compilers/accelerators, mostly • Intelligence Processors at OpenAI 🌶️

more on GPT/codex usage in Jalapeño and some of “where we’re going from there”… as article says the models really grok’d XLS, but it’s made huge strides in systemverilog as lingua franca, if you last tried before Astra it’s time for another look…
How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip spectrum.ieee.org/llms-for-c…
2
3
37
6,083
egg, and panko, on our face we know you have choices when it comes to spicy codenames and we're committed to doing better
There is a lot of deserved hype for how great OpenAI's Jalapeno ASIC and system will be. However, there is one major flaw that no one is talking about at all... (1/6)🧵
3
13
4,065
just because i'm seeing ad hominem attacks for some godforsaken reason i was the one who was scrolling and @JordanNanos is 100% right ai is what optimized it into its roofline achieving form
SemiAnalysis' Jordan Nanos reveals the OpenAI engineers scrolling their own kernel code had no idea what it did line by line, and it didn't matter because the AI does "It's not slop, because it performs, but it's literally just AI-generated assembly basically puked out in Gluon, this low-level kernel programming language that they built on top of Triton, which uses this really interesting programming model." "The point is, the guys who were scrolling through this code with us, it was kind of clear that they know a whole bunch about hardware, they know a whole bunch about the concepts in the system, and they just have no idea what this MLA kernel that they're showing us for DeepSeek actually does." "You can go line by line and it's like nope, nope, nope." "But it doesn't matter. The AI understands it, the AI tests it, and you see the results. It produces correct kernels that perform really well." "And I think this is just a sign of what's to come again." "That actual code is not necessarily something that a human has to reason about deeply, if the AI knows how to manipulate the data movement and the processing elements on the hardware that you've given it." ________ More takeaways from OpenAI's Jalapeno development: firesidealpha.substack.com/8…
9
5
55
12,604
for folks who didn't see the slide on MLA we presented at @hotchipsorg the whole idea is you can speedrun functional description to obtainable roofline all within the semantics of the machine / primitives exposed in the gluon kernel language
3
348
it is quite analogous to not knowing exactly what assembly comes out of the compiler at -O3 and why, even if you wrote the original "functional spec" of C++ code that fed the compiler
3
2
13
897
if you sat down and really studied it you could put it all back together, and the semantics are naturally limited by the kernel language, but part of the idea of ai-qua-kernel-generator is that you don't want to have to do that
6
434
How is this any different than looking at the output of a compiler and saying “I don’t know what that does line by line?” Most people can’t even read assembly code.
SemiAnalysis' Jordan Nanos reveals the OpenAI engineers scrolling their own kernel code had no idea what it did line by line, and it didn't matter because the AI does "It's not slop, because it performs, but it's literally just AI-generated assembly basically puked out in Gluon, this low-level kernel programming language that they built on top of Triton, which uses this really interesting programming model." "The point is, the guys who were scrolling through this code with us, it was kind of clear that they know a whole bunch about hardware, they know a whole bunch about the concepts in the system, and they just have no idea what this MLA kernel that they're showing us for DeepSeek actually does." "You can go line by line and it's like nope, nope, nope." "But it doesn't matter. The AI understands it, the AI tests it, and you see the results. It produces correct kernels that perform really well." "And I think this is just a sign of what's to come again." "That actual code is not necessarily something that a human has to reason about deeply, if the AI knows how to manipulate the data movement and the processing elements on the hardware that you've given it." ________ More takeaways from OpenAI's Jalapeno development: firesidealpha.substack.com/8…
3
1
15
1,790
don't forget to be excellent to one another, please...
24
1,450
great explanation style from @calebfoundry in piped.video/watch?v=yHNp_rT6… the drawings are clutch (note @calebfoundry that agentx first went live publicly on aug 10, only ~two weeks before the hot chips talk)
1
4
37
3,735
beautifully done walkthrough of the @hotchipsorg talk slides by @IanCutress /the/ place to watch until the hotchips talks are on youtube note the Gluon kernel language is the highest-performance low-level-primitives-exposed layer in the Triton ecosystem triton-lang.org/main/gluon/i…
I love the taste of jalapeños. Here's my take on the @OpenAI presentation chip that claims to be faster than @nvidia. piped.video/Ic0kYWjffjI
2
1
17
4,626
@cdleary retweeted
ai for chip design is underrated
OpenAI decided to keep the chip team small and use AI for a an improvement loop Early stages of Hardware in the loop RSI?
99
120
1,809
165,154
@cdleary retweeted
we made a chip and it is fast
4,016
2,034
48,096
8,592,020
@cdleary retweeted
We plan to begin deploying Jalapeño in OpenAI’s compute infrastructure by year-end. It’s the first step in a multigenerational roadmap: Gen 2 is deep in development, and Gen 3 is taking shape. Each generation will push efficiency and speed further. openai.com/index/jalapeno-fi…
25
26
1,082
133,816
@cdleary retweeted
inference numbers published for jalapeno, team did an amazing job openai.com/index/jalapeno-fi…
OpenAI: "Today, we shared the first measured performance results from Jalapeño, OpenAI’s first custom inference chip. On InferenceX, a public benchmark using GPT‑OSS 120B, Jalapeño delivered more peak throughput per kilowatt and lower token latency than the commercial systems in the comparison. It also performed strongly on DeepSeek R1 and Kimi K2, showing that its gains extend across model families." "This is Jevons paradox: greater efficiency makes more uses worthwhile, expanding consumption and creating new economic activity through more work completed, better decisions, more products launched, and more revenue generated." "AI played a direct role in Jalapeño’s development, enabling the team to move from initial design to tapeout in nine months by exploring implementations, shortening design, measurement, and verification loops, and continuously iterating on model workloads. AI also helped optimize the chip’s arithmetic circuits, allowing the team to fit more compute performance into the chip on schedule." "We plan to begin deploying Jalapeño within OpenAI’s compute infrastructure by the end of the year. It is the first generation of a multigenerational roadmap: Gen 2 is deep in development, and Gen 3 is taking shape. Each generation will build on what we learn and further advance both efficiency and speed." "As we prepare for deployment, we are continuing production qualification, maturing the software, preparing to operate Jalapeño at scale, and validating performance across more models."
63
92
1,343
170,262
@cdleary retweeted
OpenAI Jalapeno is spicy Usually first generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Rubin This is huge news! We got to go into OpenAI's lab to dissect their new chip and dive into the software, architecture, and performance Incredible work!
OpenAI Jalapeño: Better Than Nvidia Blackwell OpenAI’s self-designed ASIC compared with Rubin, Jalapeño’s TCO, throughput per MW, and spicy deets newsletter.semianalysis.com/…
86
121
2,131
375,725
it's a good chip ser agree w/ the article: the team is cracked and delivered like you wouldn't believe (gonna be a lot of cope) the spice must flow 🌶️
OpenAI Jalapeño: Better Than Nvidia Blackwell OpenAI’s self-designed ASIC compared with Rubin, Jalapeño’s TCO, throughput per MW, and spicy deets newsletter.semianalysis.com/…
4
7
103
10,818
@cdleary retweeted
Replying to @rohanpaul_ai
note asic does not mean “less flexible” in industry practice, the asic model is a business model: create a chip design and work with an asic platform provider to bring it to silicon cpu parts can be built this way too “asic = fixed-function toy” is a very common misconception
2
7
800