On-device ML Engineer | Reverse-engineering neural nets | Optimizing large models for the edge 💻📱

Earth
56,000+ tokens/sec at just 80 MHz. 🤯 I burned a full Transformer with KV cache into a custom chip. Designed gate by gate as a 100% digital integrated circuit. Prototyped on a FPGA. (No GPU. No CPU) Just pure digital silicon running @karpathy microGPT, spelling out names on a tiny LCD. This is GateGPT 👇
169
562
5,029
759,950
98%. That's the local win/tie rate for small models on chat - and it's a weighted average across 20+ occupational domains, not a cherry-picked slice. The "you still need the cloud for everything" assumption just got an expiration date. Huge congrats to you and the team. And getting the FT to run it? Even better 👏 Long live on-device intelligence!!!
dreams do come true 🥹. excited to see our work (w/@JonSaadFalcon, @HazyResearch, john hennessy and @Azaliamirh) feat. in a major way in @FT. the world is becoming increasingly less dependent on centralized cloud ai. we are just getting started 🚀🌖
1
1
4
495
Long live on-device intelligence. Real privacy, not brochure privacy
Today we're launching Desert Ant Labs: a European frontier AI lab building on-device intelligence. 18 models across audio, vision, and text. SDKs for Swift, Kotlin, and JavaScript. No tokens. No logins. Nothing leaves the device. desertant.com
3
183
Apple heard the community! 🗣️ Background Neural Engine access is BACK in the iOS 27 Public Beta. Huge win for on-device AI. Thanks to @argmax for fighting for it. On-device inference stays fast and efficient
Apple restricted background access to the Neural Engine on iOS 27... and it's back! This would have broken on-device inference for many apps that rely on this access, including many of our customers. We made a strong case to Apple that this access should be reinstated, channeling our customers' voice. We are elated that Apple reinstated this capability in the first iOS 27 Public Beta! Here is the step-by-step guide in our docs: app.argmaxinc.com/docs/faq#q…
8
1,146
Fabio Guzman retweeted
I am pretty impressed what open source EDA tool can achieve already. This is an programmable instrumentation amplifier. And yes. Its design and the IHP PDK used is open source :-)
7
39
314
13,116
Learning computer architecture by building a full GPU from scratch. ▪️ 16-bit ISA & assembler ▪️ High-level programming language (J++) & compiler ▪️ Branch divergence & warp scheduling ▪️ 2 compute cores Running AI model inference at 81 MHz on a 30$ FPGA 🧵
16
89
892
66,481
The tokenizer isn't plumbing I've always suspected it was quietly deciding more than we admit, and Cohere Labs just put numbers on it. Train a tokenizer on more languages than you pretrain on, and the model inherits a plasticity you can't buy back later: - ~19% higher win rates on new languages - 8x faster adaptation - near-zero cost to the primary ones The part that stays with me is that you can't retrofit it. Swap the tokenizer in after pretraining and you're still 7% behind the model that got it right at step zero. The ceiling is set before the first gradient. It's the same pattern that's been nagging me for a while: the lottery is drawn early. Not the best architecture wins, the one your first decisions could still reach. Two things I keep circling: 1. Is the tokenizer actually free, or did we just move the cost into the embedding layer where it's easier to ignore? 2. Win rate can reward sounding right over being right. A model that sounds native in Nepali isn't necessarily reasoning in Nepali. How much of that gain is fluency, and how much survives hard tasks? 
Huge congrats to @dianaabagyan who will be presenting this work next week at ACL. We asked what relatively cheap interventions like tokenizer design early on in training improve "language plasticity" of the model post-training to adapt to new languages. 🎉🔥
1
2
5
1,059
Humans didn't design this. AI just invented radio chips no engineer would ever draw. A Princeton team fused reinforcement learning, inverse design, and diffusion models to discover entirely new RF circuit topologies. Real fabricated silicon (SiGe BiCMOS): a 34–70 GHz mm-wave PA hitting 21.2 dBm and 26% PAE, plus a 100–120 GHz sub-THz PA at 12.6 dBm. Design time: weeks → hours.
AI is designing radio chips humans couldn't even imagine 👀 Researchers just developed an AI framework combining reinforcement learning, inverse design and diffusion models to automatically discover entirely new radio-frequency integrated circuit architectures. Instead of optimizing human designs, AI invents new circuit topologies, cutting design time from weeks or months to just hours. The team fabricated working AI designed chips, including a 34-70 GHz mm-wave power amplifier with 21.2 dBm peak saturated output power and 26% PAE, plus a 100-120 GHz sub-terahertz power amplifier with 12.6 dBm peak saturated output power. This breakthrough could accelerate 6G, satellite communications, radar, autonomous vehicles and next generation wireless hardware.
2
16
1,540
I burned a full Transformer into custom silicon. And Andrej Karpathy liked the project. ❤️ When someone whose work inspired your own takes a moment to notice it, every late night suddenly feels worth it. Back to building. ⚡
56,000+ tokens/sec at just 80 MHz. 🤯 I burned a full Transformer with KV cache into a custom chip. Designed gate by gate as a 100% digital integrated circuit. Prototyped on a FPGA. (No GPU. No CPU) Just pure digital silicon running @karpathy microGPT, spelling out names on a tiny LCD. This is GateGPT 👇
2
14
259
27,280
Hard agree. Agentic AI never bottlenecked on FLOPs - it's memory, routing and control flow. We over-fit silicon to dense matmul and called it progress: the hardware lottery. I burned a full Transformer into custom silicon to feel exactly where that breaks. The next 10x is co-design, not compute.
Hardware for agentic AI isn't just a compute problem. @sarahookr on why memory, routing, and complexity have rewritten what hardware needs to do, with @MilksandMatcha at @cerebras.
2
17
2,081
Built a full Transformer with KV cache in RTL, prototyped on a Virtex-5 FPGA microGPT at ~56k tokens/s. Thought this might interest you @pcuenq
56,000+ tokens/sec at just 80 MHz. 🤯 I burned a full Transformer with KV cache into a custom chip. Designed gate by gate as a 100% digital integrated circuit. Prototyped on a FPGA. (No GPU. No CPU) Just pure digital silicon running @karpathy microGPT, spelling out names on a tiny LCD. This is GateGPT 👇
3
3
53
4,301
56,000+ tokens/sec at just 80 MHz. 🤯 I burned a full Transformer with KV cache into a custom chip. Designed gate by gate as a 100% digital integrated circuit. Prototyped on a FPGA. (No GPU. No CPU) Just pure digital silicon running @karpathy microGPT, spelling out names on a tiny LCD. This is GateGPT 👇
169
562
5,029
759,950
The whole Transformer fits in 23% of the LUTs and a single Block RAM (activations + KV cache live there) But it pins 62 of 64 DSPs at 96%. The multipliers are the wall. This is the actual place-and-route on the Virtex-5
4
6
184
25,697
How it hits 56,000+ tok/s No monolithic FSM. A microcode ROM sequences modular datapath actuators - matvec, attention, RMSNorm, exp, sampler - over one true dual-port scratchpad that also holds the persistent KV cache. 1 block · 24-dim · 4 heads · ctx 16 · Q5.11, bit-exact to Python.
9
3
155
20,694
The whole Transformer fits in 23% of the LUTs and a single Block RAM (activations + KV cache live there) But it pins 62 of 64 DSPs at 96%. The multipliers are the wall. This is the actual place-and-route on the Virtex-5
4
674
Welcome VibeThinker-1.5B to MLX! 🚀 This 1.5B model is competitive with GPT-OSS-20B and MiniMax 456B on AIME2025! 🤯 ONLY 1.54GB MEMORY FOOTPRINT! ⚡️ Run it locally on your Mac now: 🤗 Model: huggingface.co/mlx-community… 💻 Github: github.com/WeiboAI/VibeThink… #MLX #AppleSilicon
11
40
286
17,389
Replying to @FGuzmanAI
Solve the derivative of sin(x²) step by step (60.93 tokens/s on an iPhone 17 Pro)
1
1,173
Running VibeThinker-1.5B on iPhone. ~1.5GB RAM usage, reasoning behavior comparable to GPT-OSS-20B. This is where edge AI is heading. huggingface.co/mlx-community…
5
16
165
10,565
Solve the derivative of sin(x²) step by step (60.93 tokens/s on an iPhone 17 Pro)
4
1,759