Xiaomi published their RL training stack for MiMo-V2.6. Reveals its cyber RL primarily targets vulnerability reproduction (using ARVO/OSS-Fuzz), not specifically exploitation. huggingface.co/datasets/Xiao…
2
2
19
2,579
Try comparing the Windows kernel to XNU 😆
Say what you will about Windows (OS), but the NT kernel really is an engineering marvel that still puts Linux to shame in many ways. The quickest way to describe it for a programmer, is NT was more like an object-oriented language, with a strong security model from day one, whereas Linux is very…not. A lot of the “good” features in Linux feel bolted-on (SELinux, Capabilities, Namespaces) because…well they were. I love to imagine an alternate history where NT won. IMO, Microsoft *should* have made an “Open NT” in the early 2000s; not fully GPL-style open, but one where a large org could say…swap out a memory allocator for their own. (they sorta did this with limited source access, but it was too restrictive) You could imagine say…an early Amazon forking OpenNT to create an “AmazonNT” for EC2, where they have a modified scheduler, network stack, whatever. But, the security+compatibility contract keeps a stable baseline on the Microsoft side. Controversial take, but if we enter this era where users are giving AI agents increasingly higher levels of access control; the Linux kernel is legitimately a poor fit. Think about it; answer the question “What exactly is this AI agent allowed to do?” On Standard Linux, it’s disgustingly messy with lots of overlap. Do you use UIDs? GIDs? ACLs? CGROUPs? Policies? SELinux? Filesystem modes? There’s not a singular coherent graph of capabilities you can point to. Too many ways you can escape an initially narrow scope. NT, by comparison, can go the route of explicitly typed resources, and then you could have these really strong centralized audit trails when an agent goes haywire…etc. I know I’m rambling, but the point is…if you were greenfielding an OS kernel from scratch, in 2026, with the intent of being forward-looking, it would *not* look like Linux. Frankly, it’d probably look a lot closer to NT, or even a BSD fork…
4
14
3,756
Comparing MiMo-V2.6-Flash in DS4 to the Xiaomi API. More work to be done on long context :) Still cool that a single M3 Ultra can compete on decode performance. Also, prefill on M5 Ultra should be much better.
4
799
And here's MiMo-V2.6-Pro-RL in DS4 on 2x M3 Ultra
3
10
6,152
Btw, there was heavy focus on accuracy for this work, so no shortcuts with kernels that have rounding errors etc. Hence, long context currently pays a penalty, but should be salvageable with some additional effort.
2
357
Tarjei Mandt retweeted
I put my @UnpromptedAU slides up at justdionysus.github.io/slide… — a bit of reflection on exploit development in the age of AI. My TL;DR is keep pushing to understand complex things, be honest with your own understanding, and use AI as a power tool to increase pace and depth.
4
68
235
31,921
I've noticed that this model (also checked the official Xiaomi endpoint) seems to have problems emitting proper JSON as well as other random quirks. How it managed to do well in benchmarks remains a mystery...
6
1,681
I've amended my open mlx-lm PR to support this model: github.com/ml-explore/mlx-lm…
Introducing Xiaomi MiMo-V2.6 — Pro & Flash. Frontier intelligence, all the modalities, built in public. 🔹 Two omnimodal models, advancing through scaled reinforcement learning 🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks 🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models 🔹 Stronger coding, computer use, 3D reasoning and creative capabilities 🔹 Open model weights, technical report, RL environments and training code Blog:mimo.xiaomi.com/mimo-v2-6
1
2
21
4,018
Pre-orders for M5 Ultra 256GB already at 16-18 weeks shipping. Hopefully this is not bad news for the 512GB version...
2
1
16
2,354
Tarjei Mandt retweeted
Can we please stop posting about rogue AI and how it’ll kill us and go back to posting about all the cool vulns it is finding?
5
15
116
8,206
Tarjei Mandt retweeted
Qwen 3.8 Flash Next support in DwarfStar is ready! @kernelpool + @ddalcu + @Spangler3000 + me combined. Prefill with this architecture is incredible! Flat up to 262K! PR here: github.com/antirez/ds4/pull/… Trying to push decode with MTP even more, and working on Q2 Model for 64GB machines as foreseen by @antirez 🚀
35
22
244
26,244
Tarjei Mandt retweeted
I only see "ExploitBench: 100%"
Full benchmarks. Absolutely insane. ARGI-AGI 3, from 7.8% to 98.6%
2
6
26
4,833
Tarjei Mandt retweeted
Exciting announcement: I am launching a new AI-focused company today named Umbriel (@Umbriel_AI)! Excited to do some cool research in the space 😎
Introducing Umbriel: a new company focused on LLM research and engineering, particularly the development of advanced AI-driven cybersecurity pipelines, as well as broader research into LLM tech. We have some exciting things planned. Watch this space! umbriel.ai/
5
17
149
47,962
Looks like a pretty big jump over Hy3! CyberGym at a respectable 78.4 (up from 51.8). 4-bit MLX quant here: huggingface.co/mlx-community…
🚀 Hy4 preview is here. 770B, 49B active, 1M context. Built for productivity. Open source frontier. Consistent affordable price. Use it. Tell us what breaks. More on Hy blog:hy.tencent.ai/research/hy4-p… HuggingFace:huggingface.co/tencent/Hy4-p… Github:github.com/Tencent-Hunyuan/H…
1
12
2,784