Cofounder/CEO, Power Dynamics (powerdynamics.ai) | Cofounder, Tech Fellowship @AspenInstitute | Senior Consultant @HBO Silicon Valley show

On a Plane
Tencent, the world’s largest gaming company, will and should do amazing things in image/video space. This vertical is basically Tencent and ByteDance’s game to lose. 🔥
🎨 Hy Image3.5 preview is live in Miora. Edits that keep what already works — same canvas, your brand rules already remembered. Free for two weeks.
3
31
3,885
Wow Xiaomi is front tier now and @_LuoFuli had the guts to live stream RL training out in the open. This is next level. So top tier from China side: - DeepSeek - Qwen - Zhipu GLM - Kimi - Tencent HY - Xiaomi Mimo - MiniMax - ByteDance Doubao 2nd tier: - Baidu ERNIE - Meituan LongCat - iFlytek SPARK - Xiaohongshu dots
Introducing Xiaomi MiMo-V2.6 — Pro & Flash. Frontier intelligence, all the modalities, built in public. 🔹 Two omnimodal models, advancing through scaled reinforcement learning 🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks 🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models 🔹 Stronger coding, computer use, 3D reasoning and creative capabilities 🔹 Open model weights, technical report, RL environments and training code Blog:mimo.xiaomi.com/mimo-v2-6
5
17
219
14,298
While Qwen and DeepSeek are teasing 2T-8T models, @TencentHunyuan is heading to a counterintuitive direction - be small but powerful. I’ve been using HY4 Preview and it’s insanely fast and good. This worths a read 👇👇
How Tencent Packed a 770B-Parameter Model into 214 GiB Shrinking Hy4 preview's weights from roughly 1.5TB to 214 GiB is one challenge. Preserving useful capabilities and practical inference speed is another. How did @TencentHunyuan tackle both? Zhihu contributor yghstill, a member of Tencent Hunyuan's quantization team, explains the engineering behind it. The parameter count remains 770B; the compression changes how those weights are represented. Four weights, five bits Sherry is the quantization algorithm, STQ1_0 the storage format, and MIX-STQ1_0 the mixed-precision allocation scheme. Each group of four weights takes values from {-d, 0, +d}, with exactly one zero. Four zero positions multiplied by eight sign combinations gives 32 possible patterns, requiring five bits. That is 1.25 bits per weight for the codes alone. Including a shared FP16 scale for every 256 weights brings STQ1_0 to 1.3125 bits per weight. The complete mixed-precision model averages about 2.38 bits per weight. Allocate precision to specific weights, not whole layers Hy4 preview has 77 MoE layers, each with 256 routed experts. The team concentrates aggressive compression on expert weights while selectively protecting other components. For the experts' gate/up projections, MIX-STQ1_0 uses IQ2_XXS on 48 sensitive layers and STQ1_0 on 29 less sensitive layers. The author reports that mixing lower and higher precision produces less error at the same average bit budget than uniformly choosing the intermediate IQ1_M format. Layer sensitivity needs more than diagonal statistics The author describes using the full Hessian, H = XXᵀ, to measure quantization sensitivity. Its off-diagonal terms capture correlations that diagonal-only imatrix scoring misses. In the team's comparison, the two sensitivity rankings had a Spearman correlation of -0.115. The chosen layers did not follow a simple “deeper means more important” rule: precision was allocated greedily by error reduction per additional byte. Fit the scale and choose the zero together This is post-training quantization, without retraining. The encoder alternates between two decisions: fitting d with weighted least squares and choosing the zero position using imatrix-weighted error. Zeroing the smallest-magnitude weight is not always the best choice. What matters is the additional weighted error introduced by making it zero. Across 1,200 rows of real expert weights, three alternating rounds reduced weighted reconstruction error by roughly 90% compared with the original ternary encoder. This measures local weight reconstruction, not end-to-end model accuracy. Compression must survive the runtime The team implemented STQ1_0 CUDA kernels in a patched llama.cpp build. In its operator comparison, STQ1_0 ran roughly as fast as IQ1_M despite the lower bit width. The author reports nearly unchanged MRCR retrieval performance and a small decline in math. Against UD-IQ1_M at a similar bit budget, the mixed-precision model led across the reported evaluations, including a gain of more than five points on MRCR. The result comes from combining compact encoding, calibrated quantization, selective precision and usable inference kernels.
2
5
34
2,661
“Pace the frontier because it’s too dangerous!!!!” “Here is Opus 5.5.”
12
1
42
2,539
Intelligence is increasingly more like capital and connections - the more you have the more you know how to access to more. The reverse is also true.
2
1
24
2,138
China’s nuclear reactors construction, nuclear power development time, capital allocation, and research progress. Sharing now so it won’t be a surprise again in a few years time.
2
6
31
3,080
If you go to Xiaomi’s flagship store in Shenzhen, it’d remind you of an Apple Store, with smartphones, laptops, tablets, headsets etc. but it also has: > cars, beautiful ones > rice cookers > electronic toothbrushes > power banks > hair dryers > vacuum robots > air purifiers > air conditioners … … All at great prices. It’s absolutely insane.
> The team estimates a 10× productivity gain, shortening the R&D cycle from one month to 2–3 days. I think Xiaomi will be a big deal. Remember: they're not an LLM startup. They make EVERYTHING. For example, my power bank. My monitor. My travel suitcase… They need industrial AGI.
4
6
74
6,563
The entire China’s AI ecosystem + handful highly impactful players in the US (NVIDIA, Hugging Face, Prime Intellect etc) are saving the world from AI Feudalism dystopian future. 🙏🫶🫶🫶
9
8
123
4,936
They split the problem cos they had to. MixRL on the tasks you can actually verify. Everything else - long-horizon, games, subjective signals - trained off to the side & merged w MOPD. Joint RL on the hard stuff just kills rollout efficiency. That’s too researchers in Chinese frontier labs turning compute restraints into a force for progress. DeepSeek did it w process rewards. Xiaomi is doing it w mixed batches + a control plane that won’t choke. When you can’t buy the cluster, you make every rollout count. Mad respect. 💪
MiMo-V2.6: The Hard Road to Scaling Up RL MiMo-V2.6 is very likely one of the largest single RL runs, by compute, that any open-source model team has undertaken to date. In an era when compute is brutally scarce, we still chose to dedicate a team of several dozen people to one goal over an extended period: scaling up RL. That takes more than research conviction. It takes a vision for AGI, respect for the unknown, and the nerve to walk straight into the hardest problems. The result is a model whose potential was built through mid-training and unlocked through heavy RL. Today, it is the number one open-source model. I strongly recommend reading the technical report. I believe it will become one of those papers that Agent RL practitioners keep reopening and discovering something new in each time. In my view, the research innovations and engineering challenges behind it surpass those of DeepSeek R1, which I was partly involved in. Some will ask: why MixRL instead of MOPD? First, they are not competing choices. We ran MixRL on verifiable tasks of moderate difficulty, including code and related agentic tasks, and found that the resulting models generalize remarkably well. Second, tasks that are difficult to verify, extremely long-horizon, or simply too challenging to include in a joint RL run are trained separately. Including them would substantially reduce rollout efficiency or introduce significant rollout staleness. We then merge the resulting capabilities through MOPD. Games, 3D tasks, and tasks with subjective evaluation signals all fall into this category. There is also a third, slightly cheeky answer. Our team is flat enough and free enough of organizational silos that MixRL simply is not difficult for us. More importantly, everyone enjoys working this way. People from different domains come together every day, driven by the pursuit of AGI and intelligence that can continuously improve itself, to confront and resolve the RL bottlenecks in each field. I will always remember the RL daily update meetings from this period. They were intense and dense, with intelligence emerging in real time. To help the open-source community focus on solving real Agentic RL problems, we have released a Qwen model distilled from MiMo RL trajectories as a stronger starting point for RL, along with 7K diverse environments and a complete RL training framework. We hope these resources will help move Agentic RL research forward. MiMo-V2.6 is only the beginning. In an era when intelligence is easy to replicate, we still choose the hard road toward self-improvement and AGI. Much of what lies ahead remains unknown. But we are willing to keep investing the time, compute, and passion required to take on one hard problem after another and work each of them all the way through, until intelligence crosses into a new regime.
1
1
16
2,124
Alibaba now announced the “most advanced chips in China” and teased for 10T model training. Once the domestic “involution”/competition is set off, that’s it. There is no going back. The machine has started running. Just watch how quickly the semi landscape changes.
DeepSeek is training a 2T model (and planning a 8T) exclusively on Huawei Ascend. The turning point is here.
9
44
414
17,657
When will people understand that regulatory by default is retrospective? To guardrail agentic AI in finance you need digital infra capable of higher throughput with cryptographically protected sovereignty. @EmanAbio - you are up! Enlighten us! 🧠
1
1
12
2,102
Average speed 130KMh. But love the confidence.
5
1
32
2,826
DeepSeek is training a 2T model (and planning a 8T) exclusively on Huawei Ascend. The turning point is here.
39
135
2,199
109,756
The Sui Basecamp stage list keeps stacking. Google DeepMind's @amoufarek. @EvanWeb3, @RaoulGMI, @kostascrypto, @jenzhuscott, @richardsocher, @BrianQuintenz. Two weeks out from Singapore 💧
5
5
50
3,721
JEV most lasted in total 2 days.
介绍比Jev快50倍,在你设备上跑的laya-mlx! 只在你的设备上占用最高1G内存 Laya是一个开源的类似于Jev的,基于文本输出概率的分类系统 我将其移植到MLX,并且做了一些性能优化! 视频中就是这个模型在我的本地M3Max上玩贪吃蛇 这个模型能够以每秒决策60次的速度玩贪吃蛇! github.com/mizorewww/laya-ml…
9
3
63
12,879
France - at least I could name Mistral if you stretched to call it forefront. What does Canada have? What did I miss?
Canadian PM Mark Carney: There are only four countries at the forefront of AI: France, Canada, China and the United States. We must join forces to ensure an appropriate framework so that AI is safe and effective.
164
8
253
132,958
OMG STFU
Professor Jiang Explains Why AI Isn’t Real “I guarantee you it's being manipulated by humans somewhere in India.” (Via Jack Neel)
Community note
Professor Jiang's claim that AI responses are manipulated by humans in India is false. LLMs like ChatGPT generate text autonomously via next-token prediction on neural nets after training; humans aid data/RLHF but not live outputs. techradar.com/computing/arti… arstechnica.com/science/2023/0…
37
9
161
17,542
Love how my friend @CodyAFriesen built this ultra analog, nerdy device for complete local AI - with Qwen, Gemma etc local models built-in, analog dials to adjust tones/styles, voice control + output. Personal data doesn’t not leave the device. Perfect for schools or kids as home/desktop tutor. Love it! 😻
Announcing the launch of Kalea! Voice first. Local, private intelligence. She is built to be a tutor, elder companion, private intelligence, and to form agentic swarms. Sophisticated, simple to use right out of the box.
4
2
17
2,515
Should I try Jev?
15
1
20
4,129
Sounds fantastic. But then he didn’t really say anything.
🚨 OBAMA ON AI: "If we are thinking about AI just in terms of how do we cure cancer or get better energy, you can do that without having agentic AI and having it just roaming free in the internet. The reason you are doing that is because you have to market a product that people will pay money for. That’s a misalignment between what our society needs and the commercial imperatives that these companies are facing, not because necessarily they’re trying to do bad things, but because they’ve got to justify these valuations. So, that’s one more reason why it is really important for us to have a competent government and a serious bipartisan conversation around this issue, and we have to do it fast. And I would encourage voters to pay attention to this. If somebody does not have a serious plan for how to deal with this, then they’re not meeting the moment, and you should probably look for somebody else."
66
14
173
16,223