Run and train models locally with the Unsloth Desktop app. 🦥 github.com/unslothai/unsloth

San Francisco, CA
Introducing Unsloth Desktop 🦥 The first desktop app to run and train models locally. • Open-source. Runs on Mac, Windows and Linux • Supports MLX, diffusion image/video, audio, GGUF • Connect Claude Code and Codex to local LLMs • 50% more accurate, self-healing tool calls + sandboxed code exec • Works for CPU + multiGPU setups - NVIDIA, AMD, Intel, Mac • Train models 2× faster with 70% less VRAM • Private web search, deep research, RAG, MCP and exports (NVFP4, GGUF) • Use Unsloth’s OpenAI-compatible API and cloud models • Securely deploy LLMs remotely and access anywhere Unsloth Desktop is now available on unsloth.ai and GitHub. GitHub: github.com/unslothai/unsloth Blog and Guide: unsloth.ai/docs/desktop
278
645
4,815
865,290
Unsloth AI retweeted
We got @UnslothAI a DGX Station! @DanielHanChen and @NaderLikeLadder checked out Unsloth’s new @Dell Pro Max with GB300 and talked about what comes next: support for more models, faster quantization, and more efficient reinforcement learning.
30
24
293
23,468
Unsloth has surpassed 500M model downloads on Hugging Face! 🦥🤗 Qwen3.8-27B GGUF is already Unsloth’s #1 most-downloaded model ever. Thanks for all your support!
51
68
1,059
47,426
Qwen-Image-2.1 can now run locally on 12GB VRAM with Unsloth GGUFs! 🖼️ The 7B model performs on par with Nano Banana 2.0. For higher quality, you can also run Dynamic FP8 on just 6GB of VRAM via offloading. GGUF: huggingface.co/unsloth/Qwen-… Guide: unsloth.ai/docs/models/qwen-…
Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨 A unified model for both generation and editing, delivering top-tier quality in a lightweight package. Highlights: 👀 - Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs. - Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images. - Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products. - Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography. Start to create your next masterpiece with Qwen-Image-2.1! 🖼️ - Blog: qwen.ai/blog?id=qwen-image-2… - GitHub: github.com/QwenLM/Qwen-Image… - Model Scope: modelscope.cn/models/Qwen/Qw… - Hugging Face: huggingface.co/Qwen/Qwen-Ima…
67
270
2,903
341,600
Qwen-Image-2.1 FP8 and GGUF quants should now run properly in Unsloth Desktop! 💜 Image gen and editing are both supported. GitHub: github.com/unslothai/unsloth
2
6
57
11,915
You can now train and run 500+ models locally with our Unsloth Docker image! 🐳 Use our new GUI or notebooks workflow. No setup required. Works on NVIDIA and AMD. Guide: unsloth.ai/docs/get-started/… GitHub: github.com/unslothai/unsloth
Introducing Unsloth Desktop 🦥 The first desktop app to run and train models locally. • Open-source. Runs on Mac, Windows and Linux • Supports MLX, diffusion image/video, audio, GGUF • Connect Claude Code and Codex to local LLMs • 50% more accurate, self-healing tool calls + sandboxed code exec • Works for CPU + multiGPU setups - NVIDIA, AMD, Intel, Mac • Train models 2× faster with 70% less VRAM • Private web search, deep research, RAG, MCP and exports (NVFP4, GGUF) • Use Unsloth’s OpenAI-compatible API and cloud models • Securely deploy LLMs remotely and access anywhere Unsloth Desktop is now available on unsloth.ai and GitHub. GitHub: github.com/unslothai/unsloth Blog and Guide: unsloth.ai/docs/desktop
23
99
708
53,928
Unsloth AI retweeted
Friendly reminder that you can fine tune 500+ open source models in a free Google Colab You can even upload PDFs/CSVs and turn them into usable synthetic datasets. 1. Open the Google Colab below 2. Run the blocks to install Unsloth Studio 3. Choose a model 4. Upload a dataset 5. You're good to go! And you can of course export your model afterwards.
7
74
591
42,991
Qwen3.8-27B Unsloth GGUF is now the #1 most-liked GGUF of all time! The model hit 10M downloads and 3.7K likes in just 24 days on Hugging Face - all thanks to you. 🤗🦥 GGUF: huggingface.co/unsloth/Qwen3… Guide: unsloth.ai/docs/models/qwen3…
80
124
1,748
92,752
Unsloth AI retweeted
ICYMI, we're celebrating 1 BILLION+ downloads for @googlegemma 💎 How are developers actually using open models? @GoogleDeepMind’s @DynamicWebPaige caught up with devs and collaborators like @UnslothAI and @Qualcomm to hear how they’re building on-device tools, running local fine-tuning, and pushing multimodal breakthroughs.
11
20
189
55,976
Unsloth AI retweeted
Local GGUF inference is now up to 3.3× faster at long context lengths
We made GLM-5.3-Flash run 3.3x faster locally! Local GGUF inference is now 1.6–3.4× faster with optimized decoding and bonus multi-token prediction. Run 3-bit on 128GB setups via Unsloth Desktop or llama.cpp. Guide: unsloth.ai/docs/models/glm-5… GGUF: huggingface.co/unsloth/GLM-5…
16
4
273
26,988
We made GLM-5.3-Flash run 3.3x faster locally! Local GGUF inference is now 1.6–3.4× faster with optimized decoding and bonus multi-token prediction. Run 3-bit on 128GB setups via Unsloth Desktop or llama.cpp. Guide: unsloth.ai/docs/models/glm-5… GGUF: huggingface.co/unsloth/GLM-5…
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: z.ai/blog/glm-5.3-flash Available now across all official platforms: Weights: huggingface.co/zai-org/GLM-5… API: docs.z.ai/guides/llm/glm-5.3… Coding Plan: z.ai/subscribe ZCode: zcode.z.ai/en Chat: chat.z.ai AutoClaw: autoclaw.z.ai
72
140
1,480
232,235
You can now run Unsloth GGUFs locally in one-click via Hermes! ✨ Qwen3.8-27B, Qwen3.8-Flash, DeepSeek-V4-Flash and more are all supported.
Hermes Desktop now sets up local models in one click. It automatically reads your hardware, picks the best model for you, then downloads it and configures the runtime.
28
75
890
73,454
Qwen3.8-Flash can now run 1.7× faster locally with MTP!⚡️ GGUFs can reach 170 tokens/s on a RTX PRO 6000. MTP enables Qwen3.8-Flash-Next ~1.3–1.7× faster inference with no accuracy change. GGUFs: huggingface.co/unsloth/Qwen3… Guide: unsloth.ai/docs/models/qwen3…
Qwen3.8-Flash can now be run locally! 🔥 The 125B MoE model outperforms Claude-Opus-4.6 (Max). Run on 75GB RAM via Unsloth GGUFs. Qwen3.8-Flash-Next enables CPU RAM / unified mem setups to deliver near VRAM speeds. Guide: unsloth.ai/docs/models/qwen3… GGUF: huggingface.co/unsloth/Qwen3…
65
129
1,190
114,841
Unsloth AI retweeted
⚡ Fine-tune and quantize AI models directly on NVIDIA Jetson. Our new Jetson AI Lab tutorial shows how to use Unsloth and memory-efficient QLoRA to customize models, export quantized GGUF files, and run them locally with llama.cpp. Follow hands-on examples for: 🔹 NVIDIA Nemotron 3.5 Lightning on Jetson AGX Thor 🔹 Qwen3.5-4B on Jetson Orin Nano Start optimizing: nvda.ws/45YagAq
13
53
381
24,000
Unsloth AI retweeted
GLM-5.3 can now be run locally! The 2-bit model retains ~81% accuracy after we shrunk it from 1.51TB to 239GB (-83% size). Run on a 256GB Mac or RAM/VRAM setups. GLM-5.3 is the strongest open model to date. Guide: unsloth.ai/docs/models/glm-5… GGUF: huggingface.co/unsloth/GLM-5…
GLM-5.3 is now open-weight. Our most capable model for agentic coding and cyber defense is now available to download, run, and customize. Weights: huggingface.co/zai-org/GLM-5… Tech blog: z.ai/blog/glm-5.3
102
212
1,898
230,051
GLM-5.3-Flash can now be run locally! ✨ Run 3-bit on 128GB RAM via Unsloth GGUF. GLM-5.3-Flash (ox-alpha) rivals Claude Opus 4.8 on DeepSWE, coding & agentic benchmarks. Guide: unsloth.ai/docs/models/glm-5… GGUF: huggingface.co/unsloth/GLM-5…
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: z.ai/blog/glm-5.3-flash Available now across all official platforms: Weights: huggingface.co/zai-org/GLM-5… API: docs.z.ai/guides/llm/glm-5.3… Coding Plan: z.ai/subscribe ZCode: zcode.z.ai/en Chat: chat.z.ai AutoClaw: autoclaw.z.ai
105
200
1,678
306,228
Unsloth AI retweeted
The VRAM barrier is officially dead. I just ran Qwen 3.8 Flash Next (MoE) 125B A6B with a 250,000 context window on a single 24GB RTX 4090. 21 tokens/sec decode. 364 t/s prefill. no mtp. no dflash. no kv cache quantization! We are running datacenter models on consumer hardware. Tested on Ubuntu 22 | CUDA 13.0 | PCIe 4.0 x16 | 110 GB DDR4 System RAM with a continuous 28k prompt across all runs. ### The Benchmarks & Scaling # 1. Hybrid Offload (-ncmoe 40 @ 80k Context) Offloaded 40 expert layers to the GPU, pushing VRAM to the ceiling. ./build/bin/llama-server -m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf -c 80000 --port 8080 -v --fit off -b 4096 -ub 4096 -ncmoe 40 Prefill: 383.85 t/s | Decode: 22.52 t/s Footprint: 23.85 GB VRAM | 97 GB RAM # 2. Full CPU MoE Offload (-cmoe @ 80k Context) Pinned all 512 expert layers to DDR4 RAM (-cmoe), keeping attention on the 4090. llama.cpp flags: (Same as above, replace -ncmoe 40 with -cmoe) Prefill: 355.72 t/s | Decode: 20.84 t/s Footprint: 11.66 GB VRAM (12GB+ VRAM freed up!) | 110 GB RAM # 3. The 180,000 Context Run Prefill: 357.75 t/s | Decode: 20.98 t/s | VRAM: 15.6 GB | RAM: 110 GB # 4. The 250,000 Context Absolute Ceiling ./build/bin/llama-server -m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf -c 250000 --port 8080 -v --fit off -b 4096 -ub 4096 -cmoe Prefill: 364.29 t/s | Decode: 20.97 t/s Footprint: 18.3 GB VRAM (Still ~5.7 GB of VRAM headroom!) | 110 GB RAM ### Key Insights: -b 4096 -ub 4096: doubles the prompt ingestion from ~150 to 364+ t/s. -cmoe Free Lunch: Shifting expert layers to DDR4 RAM slashes VRAM from 24GB to 11.6GB with virtually zero decode penalty (22.5 -> 20.9 t/s), enabling the 250k context ceiling. Qwen 3.8 Flash-Next (UD-Q4_K_XL) is a massive 111.4 GB model split across 4 shards. To run this architecture, you must build from the experimental PR branch (#27742) by @danielhanchen: git clone && cd llama.cpp git fetch origin pull/27742/head:qwen-next && git checkout qwen-next cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=native -DBUILD_SHARED_LIBS=OFF cmake --build build --config Release -j $(nproc) --target llama-server A single 4090 paired with 100 GB of cheap DDR4 RAM will comfortably serve production grade 125B inference. While Qwen 3.8 27B (dense) still holds the crown for single 3090/4090 rigs, Flash Next proves 125B hybrid models are officially viable on consumer hardware. Hugging Face GGUF link and complete performance telemetry graphs are dropped in the replies below. GLM 5.3 Flash VS Qwen 3.8 Flash Next, which one takes the open weights crown this week?
Qwen3.8-Flash can now be run locally! 🔥 The 125B MoE model outperforms Claude-Opus-4.6 (Max). Run on 75GB RAM via Unsloth GGUFs. Qwen3.8-Flash-Next enables CPU RAM / unified mem setups to deliver near VRAM speeds. Guide: unsloth.ai/docs/models/qwen3… GGUF: huggingface.co/unsloth/Qwen3…
162
456
4,710
1,019,778
Unsloth AI retweeted
A high-performance 125B model now running locally on just 75GB RAM! Thank you @UnslothAI for the day-0 support.🥳
Qwen3.8-Flash can now be run locally! 🔥 The 125B MoE model outperforms Claude-Opus-4.6 (Max). Run on 75GB RAM via Unsloth GGUFs. Qwen3.8-Flash-Next enables CPU RAM / unified mem setups to deliver near VRAM speeds. Guide: unsloth.ai/docs/models/qwen3… GGUF: huggingface.co/unsloth/Qwen3…
65
97
1,547
113,997
Qwen3.8-Flash can now be run locally! 🔥 The 125B MoE model outperforms Claude-Opus-4.6 (Max). Run on 75GB RAM via Unsloth GGUFs. Qwen3.8-Flash-Next enables CPU RAM / unified mem setups to deliver near VRAM speeds. Guide: unsloth.ai/docs/models/qwen3… GGUF: huggingface.co/unsloth/Qwen3…
⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Blog: qwen.ai/blog?id=qwen3.8-flas… - Technical Report: github.com/QwenLM/Qwen3.8-Fl… - Hugging Face: huggingface.co/Qwen/Qwen3.8-… - ModelScope: modelscope.cn/models/Qwen/Qw…
165
318
2,546
1,155,318
You can now fine-tune Qwen3.8-27B for free with our notebook! 🔥 Local training works on 24GB VRAM. Unsloth trains Qwen3.8-27B 1.5x faster with 50% less VRAM. GitHub: github.com/unslothai/unsloth Qwen3.8-27B Notebooks + Guide: unsloth.ai/docs/models/qwen3…
66
243
2,034
158,540