please codex desktop app let me have multiple tabs open :(
4
16
1,064
emilio andere retweeted
It's hard to overstate how valuable @wafer_ai's stability has been. Previous providers caused latency spikes and eval aberrations. Wafer has eliminated a class of issues we used to lose sleep over. If something goes awry, they're the most responsive partner we've worked with.
This is an awesome customer story (I’m a proud @brilliantorg investor and a huge supporter of their mission). We heard the same themes over and over again when speaking to @wafer_ai customers.
4
4
20
2,929
emilio andere retweeted
realtime inference? wafer.
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours. GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras. that's 44% lower latency with a much larger model. YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint. Wafer agents tuned the serving setup for YC’s request rate, cache usage, and prompt and response lengths. users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers. read how YC built the experience and landed on Wafer 🧵 link in thread
4
8
475
emilio andere retweeted
ultra fast inference? wafer.
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours. GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras. that's 44% lower latency with a much larger model. YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint. Wafer agents tuned the serving setup for YC’s request rate, cache usage, and prompt and response lengths. users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers. read how YC built the experience and landed on Wafer 🧵 link in thread
3
5
14
1,100
emilio andere retweeted
ultra fast inference? wafer
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours. GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras. that's 44% lower latency with a much larger model. YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint. Wafer agents tuned the serving setup for YC’s request rate, cache usage, and prompt and response lengths. users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers. read how YC built the experience and landed on Wafer 🧵 link in thread
7
6
57
4,615
realtime inference? wafer
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours. GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras. that's 44% lower latency with a much larger model. YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint. Wafer agents tuned the serving setup for YC’s request rate, cache usage, and prompt and response lengths. users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers. read how YC built the experience and landed on Wafer 🧵 link in thread
8
5
37
9,085
emilio andere retweeted
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours. GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras. that's 44% lower latency with a much larger model. YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint. Wafer agents tuned the serving setup for YC’s request rate, cache usage, and prompt and response lengths. users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers. read how YC built the experience and landed on Wafer 🧵 link in thread
27
11
98
26,179
emilio andere retweeted
Ever since we moved our tutor to Wafer, the responses feel near-instant. We didn’t expect things to get this fast so soon. Thanks to @gpuemi and the @wafer_ai team! Love using a product built by a founder who grew up on Brilliant :)
i grew up in Mexico, where it wasn’t always easy to find really high-quality educational resources that pushed me intellectually. i knew pretty early that i wanted to eventually go to one of the best universities i could in the US, and i spent a lot of time trying to get sharper on my own. @brilliantorg was a very meaningful part of that. i still remember being in early high school and absolutely spamming their logic courses. i loved them. they genuinely made me better at thinking, especially around math and logic, and helped build a lot of the skills that eventually got me where i wanted to go. so there is something very surreal about @wafer_ai now serving Brilliant :) their AI tutor, Koji, needs to respond incredibly quickly while students are working through math and coding problems. before Wafer, Brilliant was speculatively prefetching AI-generated responses so students wouldn’t have to wait. on our dedicated GLM-5.2 endpoint, they’re now getting ~250ms time to first token and 300+ output tok/s, 3x the throughput of their previous provider! that let them remove the prefetching layer entirely and cut inference costs by 50%. it’s hard to describe how gratifying it is to help power a product that helped me become a better thinker when i was a kid. very full circle moment. 🧵 Brilliant’s full story in thread.
6
16
3,064
emilio andere retweeted
Replying to @gokulr @wafer_ai
👀
5
11
1,093
emilio andere retweeted
Replying to @gpuemi
It's been such a game-changer. We used to spend a lot of effort making the ~5s response times feel fast -- multiple layers of caching, loading animations, restructuring our prompts, etc. Switching to Wafer meant we could spend that time on the core parts of the learning experience instead!
5
11
441
emilio andere retweeted
A great learning experience has a sense of flow. Having a tutor take 5-6s to respond kills that. Switching to @wafer_ai for inference made Brilliant’s tutoring sessions feel fluid and fun again, and every session metric improved along with that.
i grew up in Mexico, where it wasn’t always easy to find really high-quality educational resources that pushed me intellectually. i knew pretty early that i wanted to eventually go to one of the best universities i could in the US, and i spent a lot of time trying to get sharper on my own. @brilliantorg was a very meaningful part of that. i still remember being in early high school and absolutely spamming their logic courses. i loved them. they genuinely made me better at thinking, especially around math and logic, and helped build a lot of the skills that eventually got me where i wanted to go. so there is something very surreal about @wafer_ai now serving Brilliant :) their AI tutor, Koji, needs to respond incredibly quickly while students are working through math and coding problems. before Wafer, Brilliant was speculatively prefetching AI-generated responses so students wouldn’t have to wait. on our dedicated GLM-5.2 endpoint, they’re now getting ~250ms time to first token and 300+ output tok/s, 3x the throughput of their previous provider! that let them remove the prefetching layer entirely and cut inference costs by 50%. it’s hard to describe how gratifying it is to help power a product that helped me become a better thinker when i was a kid. very full circle moment. 🧵 Brilliant’s full story in thread.
8
5
30
2,484
emilio andere retweeted
This is an awesome customer story (I’m a proud @brilliantorg investor and a huge supporter of their mission). We heard the same themes over and over again when speaking to @wafer_ai customers.
i grew up in Mexico, where it wasn’t always easy to find really high-quality educational resources that pushed me intellectually. i knew pretty early that i wanted to eventually go to one of the best universities i could in the US, and i spent a lot of time trying to get sharper on my own. @brilliantorg was a very meaningful part of that. i still remember being in early high school and absolutely spamming their logic courses. i loved them. they genuinely made me better at thinking, especially around math and logic, and helped build a lot of the skills that eventually got me where i wanted to go. so there is something very surreal about @wafer_ai now serving Brilliant :) their AI tutor, Koji, needs to respond incredibly quickly while students are working through math and coding problems. before Wafer, Brilliant was speculatively prefetching AI-generated responses so students wouldn’t have to wait. on our dedicated GLM-5.2 endpoint, they’re now getting ~250ms time to first token and 300+ output tok/s, 3x the throughput of their previous provider! that let them remove the prefetching layer entirely and cut inference costs by 50%. it’s hard to describe how gratifying it is to help power a product that helped me become a better thinker when i was a kid. very full circle moment. 🧵 Brilliant’s full story in thread.
2
7
25
11,663
i grew up in Mexico, where it wasn’t always easy to find really high-quality educational resources that pushed me intellectually. i knew pretty early that i wanted to eventually go to one of the best universities i could in the US, and i spent a lot of time trying to get sharper on my own. @brilliantorg was a very meaningful part of that. i still remember being in early high school and absolutely spamming their logic courses. i loved them. they genuinely made me better at thinking, especially around math and logic, and helped build a lot of the skills that eventually got me where i wanted to go. so there is something very surreal about @wafer_ai now serving Brilliant :) their AI tutor, Koji, needs to respond incredibly quickly while students are working through math and coding problems. before Wafer, Brilliant was speculatively prefetching AI-generated responses so students wouldn’t have to wait. on our dedicated GLM-5.2 endpoint, they’re now getting ~250ms time to first token and 300+ output tok/s, 3x the throughput of their previous provider! that let them remove the prefetching layer entirely and cut inference costs by 50%. it’s hard to describe how gratifying it is to help power a product that helped me become a better thinker when i was a kid. very full circle moment. 🧵 Brilliant’s full story in thread.
13
14
92
19,195
RT @gpusteve: @brilliantorg cut inference costs by 50% by making its AI tutor fast enough to answer in real time instead of pre-generating…
1
13
emilio andere retweeted
@brilliantorg 3x’d their tokens per second and cut inference costs by 50%, by making its AI tutor fast enough to answer in real time instead of speculatively prefetching answers. To keep Koji, their AI tutor, responsive during math and coding lessons, the Brilliant team was speculatively prefetching AI-generated responses so students wouldn’t have to wait. On Wafer's dedicated GLM-5.2 endpoint, Brilliant reached ~250 ms time to first token and 300+ output tokens per second. This was a 3x throughput increase over their previous provider. That speed let Koji answer learners in real time, so the Brilliant team could stop pre-generating responses and remove the prefetching layer. Read how Brilliant made the switch! 🧵 Link in Thread
9
7
28
1,237
emilio andere retweeted
you'll know more about compute and memory bandwidth bottlenecks than 95% of people if you fully understand this paper this is only the fifth resource in the ai performance engineering repo btw. imagine the ball knowledge in the other ones
we launched the most comprehensive ai performance engineering repo in the world now we'll be posting every single resource follow and save to keep up with the series. links in thread 🧵 part 5: Roofline: An Insightful Visual Performance Model for Floating-Point Programs and Multicore Architectures the authors introduced a model for relating arithmetic throughput, memory bandwidth, and data reuse. applied to AI workloads, it gives a way to evaluate batching and kernel optimizations against the hardware's limits. the paper's concepts connect to these AI performance decisions: operational intensity measures floating-point work per byte of DRAM traffic after cache reuse. during dense transformer decoding, a linear layer processing a small token batch can read a large weight matrix for little arithmetic. compute and bandwidth roofs bound throughput. if weight reads limit an inference kernel, reducing transferred bytes can lower its memory-time bound. increasing peak arithmetic throughput alone leaves that bound unchanged. the ridge point marks the minimum intensity needed for peak compute throughput to be possible. batching tokens that share a weight matrix can increase work per weight byte, moving the matrix multiplication toward that threshold. computation ceilings account for limits in instruction parallelism and operation mix. for neural-network matrix multiplication, check Tensor Core use and compare throughput with the roof for the kernel's execution path and precision. bandwidth ceilings account for memory access patterns and data placement. for GPU tensor operations, changing the layout or thread-to-data mapping to coalesce scattered accesses can improve useful bandwidth. data reuse raises intensity when it reduces DRAM traffic for the same arithmetic work. a tiled matrix multiplication can reuse operands on chip; measure whether that reuse reduces DRAM byte traffic. the memory level determines which bytes to count. if a matrix multiplication reuses operands from L2, compare its L2 traffic with L2 bandwidth and its DRAM traffic with DRAM bandwidth. the authors demonstrated the model with 4 floating-point kernels on 4 multicore systems. for AI performance work, use it to assess weight reuse, memory layout, and arithmetic execution before choosing an optimization to test.
8
10
115
10,314
emilio andere retweeted
we launched the most comprehensive ai performance engineering repo in the world now we'll be posting every single resource follow and save to keep up with the series. links in thread 🧵 part 5: Roofline: An Insightful Visual Performance Model for Floating-Point Programs and Multicore Architectures the authors introduced a model for relating arithmetic throughput, memory bandwidth, and data reuse. applied to AI workloads, it gives a way to evaluate batching and kernel optimizations against the hardware's limits. the paper's concepts connect to these AI performance decisions: operational intensity measures floating-point work per byte of DRAM traffic after cache reuse. during dense transformer decoding, a linear layer processing a small token batch can read a large weight matrix for little arithmetic. compute and bandwidth roofs bound throughput. if weight reads limit an inference kernel, reducing transferred bytes can lower its memory-time bound. increasing peak arithmetic throughput alone leaves that bound unchanged. the ridge point marks the minimum intensity needed for peak compute throughput to be possible. batching tokens that share a weight matrix can increase work per weight byte, moving the matrix multiplication toward that threshold. computation ceilings account for limits in instruction parallelism and operation mix. for neural-network matrix multiplication, check Tensor Core use and compare throughput with the roof for the kernel's execution path and precision. bandwidth ceilings account for memory access patterns and data placement. for GPU tensor operations, changing the layout or thread-to-data mapping to coalesce scattered accesses can improve useful bandwidth. data reuse raises intensity when it reduces DRAM traffic for the same arithmetic work. a tiled matrix multiplication can reuse operands on chip; measure whether that reuse reduces DRAM byte traffic. the memory level determines which bytes to count. if a matrix multiplication reuses operands from L2, compare its L2 traffic with L2 bandwidth and its DRAM traffic with DRAM bandwidth. the authors demonstrated the model with 4 floating-point kernels on 4 multicore systems. for AI performance work, use it to assess weight reuse, memory layout, and arithmetic execution before choosing an optimization to test.
8
18
100
14,235
emilio andere retweeted
catch @wafer_ai on @MTSlive td :o
OPUS 5.5 ART | BERNIE SUPERINTELLIGENCE BAN | GLOBAL AI OPTIMISM nitter.net/i/broadcasts/1rGmqpRBE…
3
6
38
5,677
excited to be on soon!
We are live! Here's what we're covering today on MTS: - Discussing Bernie Sanders' new legislation calling to BAN artificial superintelligence - Redwood Research's latest analysis of Astra - OpenAI Daybreak usage in Ukraine conflict Followed by some amazing interviews: - @zck talking about data centers as 2028 politics, the historic implications of frontier lab IPOs, and more - @Noahpinion talking about the U.S.–China AI race at altitude — what winning means, chips/energy/allies, and China’s open-model export - @boazbaraktcs of @OpenAI talking about AI's impact on mathematics, some recent price improvements for GPT-6 API, and much more - @jeremiecharris + @harris_edouard of Gladstone AI talking about the coming discussion between Trump and Xi Jinping, as well as general thoughts on the geopolitics of AGI - @wordgrammer talking about recent articles on AI, the question of whether or not AI will be banned, and more @gpuemi + Steven Arellano co-founders of Wafer talking about their $40M Series A and what it takes to build AI that builds AI - @sreeramkannan + @soubhikdeb of Eigen Labs talking about Yukon beating Google’s private quantum circuit in the open, and Gemma ~2x on Apple Silicon - Rob Lefferts CVP at Microsoft Threat Protection, talking about how AI is changing the pace of cybersecurity - @sarahehunt01 President & CEO of the Rainey Center, talking about where AI’s electricity comes from and how we can work to find even more of it in the coming years - Rob Ferguson VP Tech & Strategy at Fireworks, talking about banning Anthropic Fable internally over data-access opt-in and the growing tensions between closed model APIs and proprietary data - Alex Haro CEO of Hubble Network talking about Hubble's satellite network and their latest Series C
10
714