TensorFold 0.6.1 is live.
NVIDIA RTX PRO 6000 Blackwell has landed. The 27B runs directly on it, and NVFP4 checkpoints run in their own 4-bit math, with full-precision mode when quality matters most.
Prompts that arrive together on CUDA now fill in one pass. The 27B with 8 streams runs up to 36% faster, and the slowest first token drops from 0.8 s to 0.1 s.
Flash Next on CUDA: a shared 40,900-token system prompt now answers in under a second instead of 16-56 s. Forks resume from where they split, short requests no longer queue behind long ones, and images work alongside text.
Flash Next prompts fill about 15% faster on a DGX Spark at 8k-32k.
Flash Next on Macs runs up to 10% faster at long context (64k-128k), with the same output.
pip install now works on RTX 40, RTX 50 and RTX PRO cards, with no Docker and no root.
/v1/decisions returns choices, scores and yes/no answers straight from the model.
Native Windows support is experimental.
20 community PRs landed. Thank you all!
Thank you
@MiaAI_lab @codengod @shantanugoel @AiMan_993 @edurdias @machinegenie @minviable_org and every contributor and tester.
Special thanks to
@plotarmordev @petruspennanen @volatilemarkts for running tests as well!