Head of DevRel @MiniMax_AI. Building @MiniMaxAgent and @Hailuo_AI. Make more model be open

SH/SF
We just dropped SOTA open weights for our MiniMax‑H3 video model, and so much happened in 48 hours. The community ran H3 on untested hardware ($280 gaming GPUs, fully‑offline MacBooks) and built brand‑new unplanned tooling. Extremely glad we open‑sourced the weights. Community creations thread 🧵 Hardware requirements are surprisingly modest. Users validated H3 works within 24h on: • $280 RTX 3060 (12GB) • 16‑GB VRAM cards • RTX 4090 / RTX 5090 A year ago, this‑quality video generation was only available via cloud APIs. Now a mid‑range gaming PC is your starting bar.
110
105
1,355
108,798
MiniMax Code CLI is now open sourced 🎉 and SOTA on FrontierHarness Eval After rapid iterations from day one, we’re thrilled to share this with the developer community. Grateful for every contributor and supporter
99
94
1,271
164,850
RyanLee retweeted
I believe experimental throughput and evaluation methodology are becoming the most important parts of a harness. They are also a critical step toward moving from handcrafted harnesses to auto harnesses.
Article

From Handcrafted Harnesses to Auto Harnesses: Agent Engineering Is Becoming Experimental Science

The more time we spend building agent harnesses, the more the work starts to resemble research. We still write prompts, design tools, manage context, and build recovery mechanisms. But increasingly,

17
8
44
5,623
Congrats to @genspark_ai and @FireworksAI_HQ Happy to see more companies post-training MiniMax M3 for their own verticals. MSA’s operator efficiency makes long-horizon rollouts practical, which is what lets a deep slides loop — plan, generate, self-correct — actually land in production
Meet Gen-1 Slides, Genspark’s first model designed for knowledge work. Trained with @FireworksAI_HQ from an open-weight base, Gen-1 Slides now powers Standard mode in Genspark AI Slides, delivering frontier quality decks at roughly 1/17th of Opus 5’s price. Fewer midnight fixes. Better decks by default. Read the full blog: genspark.ai/blog/gen-1-slide…
6
41
5,917
This is my favorite kind of open-source result. We released H3 at 28 steps. Then the community started optimizing it. Now, after 11,850 votes in the H3 Acceleration Arena, several accelerated versions are statistically in the top group — with the current highest point estimate coming from @larryvrh’s 6-step H3 Turbo v4. 28 → 6 steps. Not from one team, but from an ecosystem. This is exactly why we open-weight H3. ❤️ huggingface.co/spaces/multim…
26
33
591
28,438
RyanLee retweeted
🚀 Sol-H3: @MiniMax_AI H3 Video Generation Faster Than Playback 🤩 Five seconds of world. 1.653 seconds to infer. We’re releasing Sol-H3, our fastest end-to-end MiniMax-H3 inference stack yet. On one 8× NVIDIA B300 Blackwell system, it generates five seconds of 1344×768 video with stereo audio in 1.653 seconds. Across 1×, 4×, and 8× B300, Sol-H3 reaches up to a 15.54× speedup versus Base H3. Compared with 50-step Base H3 Dense on the same 8× B300 system, the four-step Sol-H3 profile delivers: • 5s: 18.250s → 1.653s (11.04×) • 10s: 50.660s → 3.732s (13.57×) • 15s: 99.513s → 6.612s (15.05×) Sol-H3 also scales across GPU counts: • 4× B300: 2.918s / 6.993s / 12.542s for 5s / 10s / 15s (12.11–15.54×) • 1× B300: 13.745s / 37.813s / 52.260s for 5s / 10s / 15s (9.45–14.29×) All figures are medians of three measured runs after one warmup at 1344×768 and 24 FPS with stereo audio. Base H3 uses 50 scheduler points (49 DiT forwards); Sol-H3 uses four DiT forwards, so this is a full-profile comparison—not an attention-only runtime change. Sol-H3 uses Dense attention on 1× B300 and SOL with INT8 QKV / FP8 output transport on 4× / 8×. Timing includes text encoding, DiT denoising, and video/audio VAE decoding; model loading, compilation warmup, and final MP4 encoding are excluded. Sol-H3 brings Sol-Engine × Sol-Attn into one full-stack runtime: • dynamic sparse attention with no retraining • fused norm, RoPE, MLP, and sparse-attention setup • fused INT8 QKV / FP8 output communication across 8 GPUs • parallel, batched VAE decoding • precomputed AdaLN caching Inside the stack: • sparse-attention setup: 1.206 → 0.285 ms (−76.4%) • VAE decode: 7.55 → 0.602 s • ~24 GB memory freed per GPU Any MiniMax-H3 few-step LoRA can plug into the same engine, and the code is deployment-friendly under Apache 2.0. For us, the bigger milestone is crossing from “fast generation” into “faster than playback.” That opens the path toward continuous 24 FPS generation and truly interactive video systems. We’re excited to partner with @reactorworld to release Sol-H3 and make it available as an API day-0. Try it now on Reactor: reactor.inc/sandbox?model=fa… 🔗 nvlabs.github.io/Sana/Sol-En… Amazing team effort—full credits in the blog. @shanasaimoe @lawrence_cjs @yitongli165665 @haopengl33 @HaochengXiUCB @songhan_mit
58
91
683
206,971
🔥 vote for best H3-acceleration version! huggingface.co/spaces/multim…
7
5
95
10,995
Impressive work
We share H3-World 🌍 The first to turn MiniMax-H3 itself into world model. No new action module. We directly convert H3’s pretrained language understanding into world control. Only 8K samples + 0.199% Trainable Params. 📄 huggingface.co/papers/2609.0… 💻 danzer1xxxxchan.github.io/H3…
5
2
68
7,397
The MiniMax H3 ecosystem is growing faster than ever! ☺️ Check out Awesome MiniMax H3 Integrations - an index tracking everything built around H3, from 24GB VRAM local ComfyUI setups to enterprise-scale SGLang & vLLM-Omni serving deployments. 🔥 What’s inside: - Hardware & Precision Maps: Quantization guides across INT8, NVFP4, and GGUF down to 8GB VRAM limits - Speed & Serving: Acceleration LoRAs, Sol-Attn, block-caching, and multi-GPU production stacks - Developer Tooling: Native agent skills, timeline directors, multi-shot motion context nodes, and h3.c for Apple Silicon Huge thanks to the global open-source community, framework maintainers, and independent developers pushing the boundaries of what’s possible with MiniMax H3! 👇 Explore the repository & contribute: github.com/MiniMax-AI/awesom…
37
41
485
44,487
RyanLee retweeted
Developers are already pushing MiniMax Code hard in the desktop app. And one request keeps coming up: bring it to the terminal. MiniMax Code CLI does exactly that. It brings the same intelligence to your terminal, so coding agents can plug directly into repositories, scripts, CI pipelines, and open-source projects—more open, more composable, and closer to how software actually gets built. The desktop app and CLI are two interfaces to the same system, powered by the same models and engineering foundations—each designed for a different way of working. Use whichever fits the work at hand, while keeping the workflow you already rely on. Get started with MiniMax Code CLI.
26
21
218
46,537
Keep open every model! This time is music -> minimax-ai.github.io/music3-…
🎵MiniMax-Music3 Next-Generation Open-Weights Production-Ready & Versatile Music Model huggingface.co/MiniMaxAI/Min…
27
15
313
16,260
Next is LLM? 😊
11
25
1,590
RyanLee retweeted
Starting today, we’re opening the MMC TUI invite-only beta 🚀 Bring it into real repos, take on tough engineering tasks, and help us shape a coding agent built for the terminal. We’ll reset usage quotas for all beta participants later this week. A huge thank you in advance to everyone joining us—we can’t wait to build this with you! If u want to join our beta testing, please DM @MiniMaxAgent
8
8
48
6,085
thanks to @xieenze_jr , now we have 3.92× on DGX Spark and 4.52× on GeForce RTX 5090
🚀 Two days later, MiniMax-H3 goes from the data center to the desktop with Sol Engine. ⚡ 3.92× faster on DGX Spark 480p · 5s · 24 FPS ⚡ 4.52× faster on RTX 5090 720p · 5s · 24 FPS Powered by full-stack kernel optimization, Sol-Attn, and cross-step caching—using the stock 33B checkpoint, with no distillation, fine-tuning, LoRA, or offline calibration. This is what agent-native optimization is built for: adapting acceleration recipes across models, hardware, and deployment settings—from GB200 to desktop AI systems. 🔗 nvlabs.github.io/Sana/Sol-En…
5
10
95
10,620
RyanLee retweeted
we are making your consumer gpu(s) on fire 🥳
MiniMax H3 😃 ( INT4, INT8, Mixed, NVFP4) Community-compiled collection of quantized and pruned weights for 12GB to 24 GB VRAMS users 👇 huggingface.co/Abiray/Minima…
11
5
283
18,321
We just wrapped up our MiniMax Beijing meetup! Had great in‑person conversations with over 150 developers — such an amazing experience. Where should we host our next meetup? Drop your city below 👇
21
6
176
9,992
RyanLee retweeted
Day-one support for Agent Plugins in MiniMax Code.
Build a plugin once and use it across compatible agent clients. Introducing Agent Plugins, an open standard developed with @awsdevelopers, @cursor_ai, @github, @code, and @vercel that packages Agent Skills and supports MCP server configurations in a shared format.
13
11
163
39,088
RyanLee retweeted
MiniMax H3: A New Open-Weight Video Model, Live in ComfyUI nitter.net/i/broadcasts/1MJgNNRPV…
7
11
83
7,738
求贤若渴 我们正在寻找真正热爱开源、熟悉 AI 视频生成技术与开发者生态的 Open Source Ecosystem Lead,负责 MiniMax 视频模型的全球开源生态建设。 欢迎转发vrfi1sk8a0.jobs.feishu.cn/in…
None of this was us. Day 1 (and the first 48 hours) belonged to the open-source community: Generate with it @ComfyUI — native support + official quantized builds, Day 0 Diffusers — the reference Python pipeline, Day 0 @deepbeepmeep — WanGP v12.41, the 5-6 GB ultra-low VRAM path, Day 1 @pipenetwork — Phosphene + MiniMax-H3-MLX (one-click Mac app + MLX engine), Day 1 @blizaine — Maestro v1.5.5 (one-click GUI + prompt enhancer), Day 2 @ModelScope2022 — DiffSynth-Studio NF4, drops floor to 7-8 GB, Day 2 Serve it @sgl_project — SGLang Diffusion, official cookbook (2×5090 / RTX 6000), Day 0 @vllm_project — vLLM-Omni, OpenAI-compatible video endpoint, Day 0 Train it @ostrisai — AI Toolkit, first trainer (T2V + I2V) 14 h after release, Day 1 @ModelScope2022 — DiffSynth-Studio training scripts (NF4 low-memory), Day 2 @kohya_tech — Musubi Tuner, third trainer (48 GB → shrinking), Day 2 the first community LoRAs running on pruned INT8, Day 2 Speed it up @nvidia SANA team — Sol Engine, 3.95× end-to-end, no quality trade-off, Day 1 sol-attn · sage-attn · EasyCache — 20-35 % each on consumer cards (and they stack), Day 1 Spectrum acceleration — ~34 % Euler sampling time in ComfyUI, Day 2 @AMD — Day 0 support on Instinct MI300X / MI355X (ROCm + SGLang) Extend what it can do Multishot workflow — chains past the 15 s limit up to 30 s with audio, Day 2 Or just run it without installing @MiniMax_AI offical API @fal · @OpenRouter · official API — hosted from launch, $0.13/s for 2K, Day 0 @wavespeed_ai · @runpod — $0.07-0.15 per clip on rented 5090s, Day 2 And everyone who benchmarked, quantized (rockerBOO NVFP4, joeygambino GGUF…), and shared numbers: @umiyuki_ai @luta_ai @onigirikila @ivanfioravanti @Spectromachina and many more 🤗 This is what open weights are for.
12
13
132
19,711