Founder, hacker, stochastic tool. @microsoft escapee. Currently @eigenlabs. Checkout my github actions emulator: notdownhub.com

United States
The Qwen 3.8 Flash Next speculative decode shootout! We can get speeds consistently above 80-90tps for contexts up to 128k, hitting over 107tps on realistic coding workloads by make more permissive speculative decoders. Here I'm comparing three different types of speculative drafters and how they impact accuracy and decode tokens per second. Speculative drafting like MTP guesses extra tokens better drafting is the biggest knob we have at increasing tokens per second. You can either build better drafters or you can build more permissive verifiers for the drafters. First we have the previously implemented typical verification (use good enough tokens), now we are adding cascade verification (how close is the drafted token to the verification's desired token) using a straight forward version (OTP) and a version that corrects for errors (TokenV3). On the first graph which shows decode tps (prefill is unchanged so uninteresting), you can see first that OTP is very fast but you're sacrificing massive amounts of accuracy. We don't really want to do that, we want to pick the two that preserve accuracy. On the second graph which shows HumanEval+ accuracy (left bar) and task completion* (right bar), you can see here that for our tests, TokenV3 with the 0.95 setting and Typical Verification with the 0.2 setting bot have a ~20+% increase over exact verification with TokenV3 slightly being faster than Typical 0.2. So you can see that we can get increased performance with less exact verifiers without losing much accuracy. *task completion is the percent of tasks didn't infinite loop github.com/youssofal/MTPLX/p… github.com/youssofal/MTPLX/p…
3
2
26
2,605
TheDavidTai retweeted
Xing4.0 ✨ from China Telecom has been quietly trending on the Hub since last week - 29B/4B MoE - Trained entirely on Ascend 910C with MindSpore 👀 - 256K context (up to 512K) - Apache 2.0
3
4
64
4,094
Well, Polymorf just destroyed the 166tps OMLX reference benchmark I was using less than 24hrs after launch on mlx.fast. 150tps->Almost 200tps in one step 0_o
4
3
29
1,137
MLX.fast is launching a 1 week contest to improve Bonsai 2 from @PrismML! Bonsai 2 is a near losslessly comrpessed version of Qwen 3.8 27b that is around 10% the size of the original. It can fit on 16gb Apple devices and even some phones. This is a model almost anyone can run so if you weren't able to participate in the other contests, this one is for you! The community has been able to double the speed of Bonsai in under 12 hours so there is alot of headroom here. This contest runs until October 1, have fun!
8
15
71
10,222
Great Job with this!
Got temporary access to a Mac Studio M5 Ultra. Barely optimized, Qwen 3.8 Flash Next runs at 85.9 tok/s and 6085 tok/s prefill on long prompts. A 2x DGX Spark recipe gets 52.1 tok/s single stream and about 2960 tok/s prefill on 16k–64k prompts. Already ahead of 2x Spark, but I was expecting a much much better baseline…
1
4
903
Philosophers argue that human + tech merge to gain degrees of immortality has been a thing ever since writing was invented.
would you merge with the machines in order to stay alive?
125
TheDavidTai retweeted
The legend of super intelligence! @JensenHuang
5
13
173
10,198
It sounds like people are sandboxing their clankers in the super secure container platform known as k8s.

ALT Dog Cant Contain GIF

I keep telling you guys, learn fucking k8s, learn a little ML, get billions of jobs
6
777
TheDavidTai retweeted
We have been very excited to see how the community has been utilizing our Bonsai 2 model, well beyond use cases that we had originally envisioned it for. We’ve spent the last few days putting Bonsai 2 27B through various examples inspired by what the community has shown us. These examples showcase the strengths, and some of the shortcomings–particularly on longer multi-turn agentic workflows, and will be extremely helpful as we continue to improve. We’ll be sharing a few of those demos, along with some practical guidance on the settings and prompting that get the best results from the model. Demo repo: github.com/PrismML-Eng/Bonsa…
14
10
133
14,479
This is one of the biggest endorsements of Local LLM by a major US lab!
Run open models like Gemma 4 completely offline in the Antigravity SDK. Built on Google AI Edge’s LiteRT, you can now run Gemma 4 directly on your local GPU. Zero API costs, total data privacy, and no internet required.
3
1
26
2,219
cheapo inference and other stateless devices make a whole lot of sense compared to nvidia gpus
Project Suncatcher is launching a prototype datacenter satellite October 1st on a SpaceX Falcon 9. Some new info in this NYT report: plans for a flock of more than 80 networked satellites flying in formation, and a custom satellite the length of a soccer field.
1
4
299
TheDavidTai retweeted
Project Suncatcher is launching a prototype datacenter satellite October 1st on a SpaceX Falcon 9. Some new info in this NYT report: plans for a flock of more than 80 networked satellites flying in formation, and a custom satellite the length of a soccer field.
46
56
543
47,152
We are so back, crypto tokens are worth spending ai tokens to hack now.
JUST IN: Bitget crypto exchange reportedly hacked with over $170,000,000 stolen.
3
220
TheDavidTai retweeted
guys literally only want one thing and it’s fucking disgusting
Ranju
372
144
4,757
1,687,935
TheDavidTai retweeted
Treasury Secretary Scott Bessent: The US needs more open-source AI models. "We can't let these large labs have regulatory capture because that will stop innovation." He also said that stronger US open models are a strategic weapon against China, whose models are heavily built by distilling capabilities from American models. --- (full video on 'GOPFinancialServices' YT channel, link in comment)
40
86
432
40,998
TheDavidTai retweeted
The PC is about to become a lot more personal. Today, I’m excited to announce a partnership between @AMD and @perplexity_ai to bring Perplexity Portable Computer to AMD Ryzen AI Halo. This is more than running AI locally. We’re creating a new kind of agentic PC that can understand your work, reason through problems, and take action on your behalf, right on the device where your work happens. With one click, users can bring powerful AI models and agents to their PC, work with local files and applications, and run recurring workflows on device. When a task needs more advanced research or reasoning, Portable can tap into the cloud with permission. Your models. Your agents. Your data. Your stack. For decades, the PC has been a tool for getting things done. Personal AI changes that. The PC can now become a true partner in the work, helping turn information into insight and ideas into action. Big thank you to @AravSrinivas and the entire Perplexity team for making this possible. The PC transformed how we work and create. Now we’re creating the platform for the era of Personal AI. And this is just the beginning.
12
47
506
137,163
Only abundance for me but not for thee
1
175
TheDavidTai retweeted
Today, we release an experimental DSpark draft model for LFM2.5-VL-3B, bringing speculative decoding to our vision-language models. A lightweight drafter proposes multiple tokens ahead, and the target model verifies them together in a single pass. This accelerates generation without changing output quality. Across six vision-language task categories at batch size 1 and temperature 0: > MLX on M5 Max: up to 3.13x faster decoding and 2.62x end to end > llama.cpp on M3 Ultra: up to 2.14x faster decoding and 1.77x end to end > SGLang on H100: up to 2.66x faster decoding and 2.27x end to end All evaluations were collected using Pipette, the benchmarking infrastructure behind Liquid AI's public device-performance data. 🧵
23
62
474
50,518
Someone create a program that lets you extract patterns in your size for a dress you make and also a list of materials. Thank me when you become a millionaire!
生地選びから裁断・裁縫までできちゃう本格派ドレスクラフトシム『Dressmaker』Steamにてリリース!ウィッシュリスト40万件超の注目作 gamespark.jp/article/2026/09… 依頼品を納品するのみならず自由にデザインしたドレスを店頭に並べて販売することもできます。
23
368
7,880
176,223