Abdul Bari Khan retweeted
Decision 2.0 is out for vLLM Semantic Router! Open models from 0.6B to 27B. Answer 64 questions about a request in one forward pass. Apache-2.0 licensed. Loads with Transformers. Great work by Xunzhuo Liu and the team! lnkd.in/p/guwfsr5K
20
70
505
28,925
GrokBot is an underdog. Not many are talking about it after the release of OpenAI dots and Muse by Meta. They all are doing basically the same thing. GrokBot has been doing it for a long time. Offering compute for free of cost is the way forward.
19
Grok 4.7 is a major update, much better than all the previous releases. Hope to see 4.8 be a huge leap. #Grok
25
New app by @elonmusk is cooking to rival #Instagram
JUST IN: Elon Musk declares, "Instagram is for girls."
19
Must read
You do not pick a model and a GPU and call it done. You pick a file encoding, and a kernel path. The GPU follows those. Start with the loop. Text becomes tokens. Tokens move through a Transformer. Attention decides which earlier tokens matter. The runtime keeps a KV cache so the model does not recompute the whole conversation every time. Then it picks the next token and does it again. The model is not writing a whole answer in one shot. It is generating one token at a time. That loop has two phases, and they are not the same job. Prefill reads the prompt and builds the first KV cache. It is compute-heavy. That is the pause before the first word. Decode writes the answer one token at a time. It keeps rereading weights and the cache. It is bandwidth-heavy. So an RTX PRO 6000 with 1.8TB/s beats the DGX Spark with 273GB/s. That is the typing effect you feel. Long prompts punish prefill. Long answers punish decode. Long chats punish both, because the working memory grows. The inference engine is not the model. It is the traffic cop, the memory manager, the kernel dispatcher, the scheduler, the cache accountant, and the API surface. It loads the weights. It tokenizes the input. It runs the forward pass. It samples the next token. It keeps the KV cache. It streams the result. Serious engines also pick kernels. A kernel is not "the model." A kernel is a specific tensor program: shapes, layouts, datatypes, and what the silicon is actually allowed to multiply. Same math on paper. Different kernel. The file format matters. It decides what can load, what can quantize, and how fast it runs. Quantization is not one switch. Storing weights in 4-bit is not the same as doing 4-bit math. Weight quantization shrinks the model. The live context is a different thing. A Q4 sticker is not universal. The right format is the one your engine has optimized kernels for. Assuming every quantization label is portable is how people buy a 5090, download "NVFP4," and still miss out on performance. Here is the worked example. I put Qwen 3.8 27B on an RTX 5090. Two downloads. Same model. Same GPU. Both folders said NVFP4. One of them does 4-bit math for real. Weights in 4-bit. Activations in 4-bit. The matrix unit can multiply them as 4-bit times 4-bit. The other stores 4-bit weights. Then unpacks them in the kernel. Then does 16-bit math. Different kernel path. The 5090 did not choose that. The checkpoint did. Especially whether the activations are 4-bit too. If they aren't, there is no legal 4-bit times 4-bit multiply to run. The engine falls back. Quietly. The file still loads. That is the difference between an encoding you can load, an operation a backend implements, and arithmetic the hardware actually executes. They not interchangeable. This is also why prefill and decode do not get the same gift. Prefill has enough token rows to keep the matrix units busy. Native 4-bit math can matter there. Decode is still walking the weights and the cache, one token at a time. If the live state never went 4-bit, you do not get a 4-bit win on that part. The bottleneck moves. Do not benchmark "the model." Benchmark the stack you will actually run. Separate prefill from decode. One more thing people get wrong with the word Blackwell. It is a marketing name on four chips that cannot run each other's kernels. A data-center B200 is not a 5090. A 5090 is not a Spark. A Spark is not a Thor. Instruction support is necessary. It is not sufficient. A checkmark on the spec sheet is not the kernel that ran. In the inference engines article under my profile, I asked a question I still want on the wall: What quantization format has optimized kernels on my target engine? This Qwen run is that question in a box. Same GPU. Same NVFP4 label. Different kernels. The engine followed. The checkpoint decided which kernel it was allowed to launch. The file loaded but that is NOT the same as the kernel you wanted for the GPU you bought.
15
Expected. Finally they will learn from Apple release.
Don't blame me for constantly complaining about Samsung's software; this is what I go through with my Fold8 every day it's a painful experience.
51
Security, without the slideware. Apple wins device isolation. One OEM, Secure Element, Secure Enclave, per-txn cryptogram. UPI’s crypto/rail is fine. People lose money to fake QR, phishing, and mandates — not because NPCI’s switch is “weak.” Samsung is close to Apple on cards, with a larger Android attack surface. #ApplePay #AxisBank #India #UPI
32
Abdul Bari Khan retweeted
JUST IN: 🇺🇸🇮🇶 US officially withdraws all its forces from Iraq, ending 23 years of military involvement.
334
2,376
20,807
547,914
Apple should open doors to have multi region payment option. Not locked to the region ID belongs to.
Say hello to Apple Pay in India! 🇮🇳 We’re so excited to bring Apple Pay to India today, and make it easier than ever for our customers to make safe and secure purchases on Apple devices.
31
Finally. Butter late then never.
Apple Pay Reportedly Launching in India Today macrumors.com/2026/09/29/app…
41
Planning for an always-on, agent-only machine. I'm also looking to onboard #Hermes. It’s a bare-bones setup, and I'm considering a Proxmox cluster node. Any setup suggestions? #GMKtec
2
1
379
Thank you for the discount.
25% discount ends in 1 hour mia-ai.net/book
15
Abdul Bari Khan retweeted
JUST IN: 🇺🇳 OpenAI agents bombarded UN website with requests using aggressive techniques to access data on the system, WSJ reports.
127
554
4,046
178,675
#chatgpt #openai finally catching up. @grok your thoughts. Any article on detailed comparison we have ?
Get ready.
1
25
Abdul Bari Khan retweeted
MiniMax's latest text model, M3.1-Flash-Preview, debuts today on MiniMax Code. Built for everyday development, it's fast, reliable, and ready for real work, from quick bug fixes to full features.
141
199
2,222
326,335
Tested TensorFold by @ashhart on one DGX Spark: Qwen3.8-27B + DFlash2 drafts, claims vs our own run. ✅ Its own speed reproduces: 51 tok/s vs 49.6 claimed
✅ Drafted = serial output, SHA-256 identical 3/3
⚠️ The "~3× vs vLLM" is closer to 2.1–2.7× #DGXSpark #LocalLLM 🧵
TensorFold Inference Engine is here 🚀 I spent six months making one weight read count for more than one token on Apple Silicon. Draft tokens run through parallel lanes; the model verifies them together and keeps only what passes. Qwen 3.8 27B MLX 4Bit - 120-124tks Nemotron Lightning MLX 4Bit - 188-206tks Qwen3.8 Flash Next MLX 4Bit - 88-92tks CUDA Implementation is in Alpha showing strong gains. The Repo is in the comments 👇🏼
2
2
435
The exactness claim is the impressive part: same request drafted vs "draft": false gave byte-identical output (greedy code, seeded T=1 prose, JSON). Plain decoding is 13.6 tok/s, so drafting is 2.3–5.6× on the same engine.
2
24
Replying to @vllm_project
Quality: GSM8K (20 problems, thinking on) TensorFold 19/20 · llama.cpp 19/20 · vLLM 18/20 No accuracy cost at this sample size. Peak SoC temp 88–90 °C for all three.
33