speedmaxxing @cerebras prev @code

San Francisco
That DHH saying out loud that the age of writing code by hand is over for the industry, and causing such a huge uproar is a bit interesting, given this was clear enough since ~January for most of us using the models + paying attention. Removed paywall: newsletter.pragmaticengineer…
147
220
2,499
162,283
Software engineering was never about typing code When I was on the VS Code team, engineers often drove new product features. The backpressure came from demoing our work live every day in standup. New features had to solve a real problem, fit into the product, and improve developers’ workflows. If you failed to tell a coherent product story for your feature, the code didn’t ship. Nothing trains product judgment like being roasted by your team every day. I’ve spent the last year doubling down on everything around the code. Thanks @kentcdodds for a great conversation on the evolution of product engineering in the age of AI!
Five minutes. List the paper cuts. @joyceerhl on the friction log that builds product taste.
8
6
82
15,280
Five minutes. List the paper cuts. @joyceerhl on the friction log that builds product taste.
2
2
24
24,316
Joyce retweeted
you can just hallucinate the entire internet with Qwen 3.8 27b running at 2,000 tokens/second? part 2 of turning @cerebras + @Alibaba_Qwen 3.8 27B into an OS: built an offline browser with zero network calls and mounted it directly the JIT ubuntu desktop. no wifi. no scraping. zero packets sent to external CDNs. you search a site, set a year, and qwen 27b at 1,950 tok/s synthesizes the entire DOM on the fly. here is youtube in 2045 vs 1999: → search google for youtube inside the OS → scrub to 2045: instant futuristic feed → scrub to 1999: raw web 1.0 time capsule in seconds at this speed, browsing isn't retrieving files from a server, it's querying an alternate reality. a 2D browser window is just step zero. imagine full operating systems, virtual worlds, and complex simulation engines existing purely as model weights. zero gigabytes stored on disk, just pure interactive reality streamed on demand. What else becomes a possibility with the qwen 3.8 27b (dense) at 2000 tokens/sec?
Operating System powered by Qwen 3.8 27B at 1950 tokens/sec! here is what 1,950 tokens/second @Alibaba_Qwen's 3.8 27b actually looks like on @cerebras: i wrote a minimal python web server that turns cerebras inference into a live operating system. zero apps on disk. when you double click an icon: 1) python proxies a raw SSE stream from qwen 27b at 1,950 tok/s 2) calculator compiles & mounts in 11s 3) full canvas paint studio with brush engine compiles in 10s. at 2,000 tokens/second, software is just an on demand hallucination that runs instantly. the model weights ARE the operating system runtime. what else would you build at 1,950 tokens/second?
159
353
4,980
1,666,703
happy @code documentary day to all who celebrate 🫶
🍿 The Story of VS Code premieres today! 🔔 Join us for the world premiere at 8:00 AM PT: aka.ms/the-story-of-vs-code piped.video/watch?v=kHL3Xzjp…
1
1
26
9,673
Joyce retweeted
Ten years ago, we started @cerebras around an approach many believed was impossible. As a computer architect, it is hard for me to imagine a more exciting time. Model releases are accelerating, and hardware tapeout is compressing from multi-year roadmaps to annual launches. Hot Chips is my favorite conference, and it’s where I launched Cerebras 7 years ago. This year’s conference was especially exciting, and so much innovation was shared. I am watching the industry recreate itself: SRAM is mainstream, DRAM is moving into the third dimension, networks are being fundamentally redesigned, and AI is helping design and program the chips themselves. The industry has never moved faster and some of the hardest architectural questions are still wide open.
74
157
1,733
685,237
Joyce retweeted
The Fastest AI Just Got Faster. Meet CS-4.
165
358
4,354
1,524,766
Joyce retweeted
Ultrafast mode for GPT-5.6 Sol is now in limited preview, powered by Cerebras. We gave @OpenAI's GPT-5.6 Sol the same prompt on Ultrafast and Standard: build a financial terminal-style dashboard for analysts. Ultrafast: 1 min 50 seconds Standard: 12 min 20 seconds Same result, nearly 7x faster.
73
122
1,962
145,770
We're excited to preview Ultrafast mode: GPT-5.6 Sol at up to 14X the speed. It has been fun partnering with the @cerebras team to make it possible to achieve speed without compromising intelligence! We've been using it internally, along with select enterprise customers, and the results are incredible If you want to get notified when we expand access, sign up here: openai.com/form/ultrafast/
9
7
100
7,735
Joyce retweeted
Previewing Ultrafast mode for @OpenAI's GPT 5.6 Sol, powered by Cerebras. GPT-5.6 Sol Ultrafast generates responses at up to 750 tokens per second. That's the full, GPT-5.6-Sol model -- up to 14× faster than the same model on Standard processing. It speedran Humanity’s Last Exam in 11h 11m, nearly 7× faster than Claude Fable 5 with comparable accuracy.
157
363
4,447
1,289,920
Joyce retweeted
Previewing Ultrafast mode: GPT-5.6 Sol at up to 14x the speed. Launching first in the OpenAI API to a select group of customers with expanded access to more businesses as capacity grows.
795
957
14,943
4,136,468
was amazing chatting with Sarah and the @cerebras team! (visceral excitement about a future of ultra-fast, ultra-intelligent inference)
“Whatever hardware allows that to happen is the hardware that we want.” @jeffreygwang from @OpenAI on a future where agents consume far more tokens. Cerebras is building the infrastructure to deliver those tokens at speed.
1
4
48
6,296
Joyce retweeted
Today, @AMD and Cerebras introduced a powerful disaggregated inference solution, pairing the right engine to each phase of the inference pipeline. This is what agentic AI has been waiting for: the fastest production inference at massive scale.
51
222
2,248
666,416
Joyce retweeted
Best hackathon prize yet! @cerebras Thanks for the Mac Mini and glad you liked what I was working on !
2
3
12
1,979
Joyce retweeted
At 1000 tps this is the most intelligence per second you can get today. Wouldn’t be possible without our partnership with @cerebras
Introducing SWE-1.7, the most capable model we’ve trained yet. It scores within a few points of the strongest frontier models at a fraction of the cost, and is now available at 1000 tok/s. RL is not hitting its limit: after refining our recipe, we keep seeing gains as we scale
1
6
44
3,214
Joyce retweeted
generative user interfaces at the speed of thought. you can now build "tab autocomplete" for every app. ultra-fast inference @cerebras & your components render by the @tambo_ai agent.
11
16
219
24,458
Joyce retweeted
We’ve made GPT-5.3-Codex-Spark about 30% faster. It is now serving at over 1200 tokens per second. More to come on speed across the board.
207
116
2,553
350,740
Joyce retweeted
GLM4.7 on @cerebras is insane. I was impressed with the model performance on my sample test, but I was BLOWN AWAY by the speed. Real window into the future of AI.
13
15
218
63,936