The TensorFold engine lands on the Sparks 🫡
Showing roughly 2x gain across the board with headroom for more.
Point your agent at the repo, enjoy the speed 🚀🚀🚀
tensorfold.dev@NVIDIAAI@NaderLikeLadder@sundeep
TensorFold Inference Engine is here 🚀
I spent six months making one weight read count for more than one token on Apple Silicon.
Draft tokens run through parallel lanes; the model verifies them together and keeps only what passes.
Qwen 3.8 27B MLX 4Bit - 120-124tks
Nemotron Lightning MLX 4Bit - 188-206tks
Qwen3.8 Flash Next MLX 4Bit - 88-92tks
CUDA Implementation is in Alpha showing strong gains.
The Repo is in the comments 👇🏼
The need for INSANE speed? Yes ⚡️
DeepSeep v4.1 Flash for 4x DGX Sparks is getting... ridicules 🤯
Decode improvements
- Prose decode is now 87 tok/s single stream
- On 4 streams it's 163.6 tok/s
- Code decode is 124.8 tok/s single stream
- On 4 streams it's 246.8 tok/s
Prefill numbers?
5800-5925 tok/s on 16k-128k
5394 tok/s on 256k
It feels like a FAST API! If you have 4x DGX Sparks, you need to try this out ASAP. I wonder if the M5 Ultra 512 GB could beat this.
The improvements for now are only for TP=4 and do not effect TP=3. Thanks to @majewskizby for this awesome PR!
Get it here:
github.com/MiaAI-Lab/DeepSee…