Nanbeige LLM Lab, BOSS Zhipin (Kanzhun Limited)

Beijing
Pinned Tweet
We released Nanbeige4.2-3B, its Looped Transformer increases model capacity without adding parameters, delivering a capable 3B agent. Nanbeige4.5 is training with LoopSplit, mHC+depth attention & concatenated n-gram embeddings,already in the modeling code. huggingface.co/Nanbeige/Nanb…
37
125
926
214,737
With the release of OpenAI's GPT-6 Astra @openai and @rasbt's widely shared article, Nanbeige4.2-3B ( huggingface.co/Nanbeige/Nanb…) is back in the spotlight! 🚀 Built on a Looped Transformer, our 3B model consistently outperforms larger models (e.g., Qwen3.5-9B, Gemma4-12B) across key Agent benchmarks: SWE-Bench Pro, Terminal-Bench, GDPval, and more. Nanbeige4.5 is currently training with LoopSplit, mHC + depth attention, and concatenated n-gram embeddings. Stay tuned!
A lot of hype around OpenAI's Astra model here on my timeline today. Apparently, this goes back to a new article from The Information, which said Astra is a "recurrent depth or looped transformer". It's always interesting to read about new or different approaches (including rumors about what the closed labs may be up to), but let's debunk this a bit. About 2 months ago, I shared the architecture details of Nanbeige, for example, where "Nanbeige4.2-3B is pretrained from scratch on 28T tokens with a Looped Transformer that reuses the layer stack to increase capacity without adding parameters." Yes, that's it. The looped transformer idea is just reusing layers in the transformer block. In the case of Nanbeige, the main idea is to reuse the same 22-layer stack (=transformer block) twice instead of once. So, effectively it extends the 22-layer architecture to 44 layers, but without duplicating the weights. In simple terms, this roughly doubles the size of the model (if we ignore the embedding and output layers for a second). But instead of requiring 2x the storage and RAM to host this model, it stays at the same size since we reuse the components. However, it's almost 2x as expensive in terms of compute, because we run the embedded text through almost 2x as many layers. Why? In the Nanbeige 4.2 technical report, the researchers found that two passes gave the best trade-off and retained about 75% of the token efficiency of a standard architecture. (More passes gave barely any gains but made the training much slower and much more expensive.) While, as far as I know, Nanbeige 4.2 is the first notable open-weight model that adopted this approach, the idea goes back to the NeurIPS paper "Mixture-of-recursions: Learning dynamic recursive depths for adaptive token-level computation". Actually, this paper proposes a mechanism that is a bit more sophisticated by adding a learned router that determines whether each token receives one, two, or more passes. So, easy tokens can exit early while harder tokens receive additional computation. In sum, Astra may be a really good model, but this shouldn't be about this "looped transformer aspect," which is just a tiny architectural tweak. Also, the statement "the new technique works in a way that obscures some or all of the AI's reasoning, otherwise known as 'chain-of-thought'" is not necessarily true with respect to the looped transformer method. It's possible that The Information journalist refers to some other technique or misunderstood the looped transformer method. Reusing layers does not by itself suppress visible chain of thought. It adds computation in hidden states before the next token is emitted, just as ordinary transformer layers do. But based on the information we have, the only plausible interpretation here is that if a model uses more of these recurrent passes, it may need to generate fewer intermediate reasoning tokens. So then more of its computation happens in latent activations that cannot be read as text. But we would get the same effect if we were scaling up the model size, like GPT 5.6 Luna -> GPT 5.6 Sol.
16
46
377
24,754
We just released the DSpark weights for Nanbeige4.2-3B (huggingface.co/Nanbeige/Nanb…).
Osaurus is coming to Windows. Same idea. Your agents, your files, your models, on your machine. This is a 3B model reading my spreadsheets and finding a margin problem I missed.
4
12
85
8,852
Artificial Analysis @ArtificialAnlys just dropped their new mobile device benchmarks, and Nanbeige 4.2-3B ranks TOP 1 in Average Score (16K max context).
Replying to @ArtificialAnlys
At the standard 16K context limit, Nanbeige4.2-3B (Reasoning) and LFM2.5-2.6B (Reasoning) tie for the top average evaluation score at 63, ahead of Ornith-1.0-9B (Reasoning, 62), Qwen3.5 9B (Reasoning, 61), Ornith-1.5-9B (Reasoning, 61), Gemma 4 E4B (Reasoning, 60) and Qwen3.5 9B (Non-reasoning, 60). Ornith-1.5-9B slips behind its older 1.0 sibling purely on a weaker instruction following performance in IFBench When the context limit is raised to 64K (represented by dots in the image), Ling 3.0 Tiny takes the top spot at 66, followed by Nanbeige4.2-3B at 65, and both Qwen3.5 9B (Reasoning) at 64 and LFM2.5-2.6B at 64
14
14
242
21,039
Nanbeige retweeted
Yes, open-source / open-weight models are important for a healthy AI ecosystem. That's how we can verify things, check claims, and keep up outside the closed labs. Plus, it gives us the freedom to run AI on our own hardware if we are not ready to share personal data and IPs with closed labs through using their models. (Not that proprietary models are bad, actually I use them a lot as well, but it wouldn't healthy not to have any alternatives.) Anyway, while pretty much everyone is waiting for the Kimi K3 and Ling 3.0 weights to land on the model hub any day now, there were quite a few other interesting new open-weight model releases the past week. Yes, one of those weeks! So, here are the architecture pics along with some notes on what I found most interesting: 1) Nanbeige 4.2 3B uses looped depth sharing. This basically means it runs the same 22-layer (=transformer block) stack twice. So, it extends the 22-layer architecture to 44-layers, but without duplicating the weights. (2x the transformer block compute but same memory footprint.) Why? The info is a bit sparse, but section 2.1 of the Nanbeige 4.2 technical report says two passes gave the best trade-off and retained about 75% of the token efficiency of a standard architecture. More passes gave barely any gains but made the training much slower and much more expensive. 2) Laguna S 2.1 is poolside's Laguna model in a really nice size: 118B sparse MoE with 8B active parameters and a 1M-token context window. Otherwise, the architecture is pretty standard. It uses 36 sliding-window and 12 global (gated-)GQA layers. However, given this size, and the fact that it (just barely) runs on my DGX Spark (uses about <80 GB of RAM), this is right now the most interesting model for me personally. It's 3x bigger and thus a tad slower but maybe a good candidate as daily-driver-Qwen3.6-35B-replacement. (Still waiting on some more independent performance benchmarks though.) 3) Motif-3-Beta is a new 314B-A13B sparse MoE that is somewhat based on DeepSeek V4 in terms of mHC and latent attention. But it uses a new component, Grouped Differential Latent Attention, which is inspired by Multi-head Latent Attention. I probably should write an article about this some time, but for now, the tl;dr is as follows. Regular MLA compresses the keys and values into a smaller latent representation to mainly reduce the KV cache size. GDLA does a similar low-rank compression but puts the attention heads into groups and also learns a noise head for each group where the noise gets subtracted for filtering purposes... Anyway, a topic for another day! 4) Solar Open 2 is a new 250B-A15B hybrid MoE by Upstage that interleaves three Kimi Delta Attention layers with one GQA layer. 5) Antares 1B is a small model (and there is also an even smaller 0.3B variant) from Cisco starts that with the IBM Granite 4.0 1B backbone and uses SFT plus GRPO for terminal-based cybersecurity stuff. It is a nice example of task-specific post-training on a genuinely small model. 6) BTL-3 is a rank-32 LoRA adapter for Qwen3.6-27B aimed at coding agents and structured tool use. The really strong benchmark performance suggests that LoRA adapters are still a useful tool/technique in 2026. I added all six to the LLM Architecture Gallery for some additional details: sebastianraschka.com/llm-arc…
93
256
1,605
75,047
We’re excited by how far a 3B model can go—and even more excited about what comes next for compact local agents. @openclaw
We released Nanbeige4.2-3B, its Looped Transformer increases model capacity without adding parameters, delivering a capable 3B agent. Nanbeige4.5 is training with LoopSplit, mHC+depth attention & concatenated n-gram embeddings,already in the modeling code. huggingface.co/Nanbeige/Nanb…
12
43
425
41,063
Nanbeige4.2-3B technical report is out on @huggingface ! huggingface.co/Nanbeige/Nanb…
We released Nanbeige4.2-3B, its Looped Transformer increases model capacity without adding parameters, delivering a capable 3B agent. Nanbeige4.5 is training with LoopSplit, mHC+depth attention & concatenated n-gram embeddings,already in the modeling code. huggingface.co/Nanbeige/Nanb…
6
21
175
40,074
Nanbeige retweeted
A 2.4GB model just went toe-to-toe with today’s best local LLMs… and almost beat a 26B MoE. On an RDNA1 GPU. Nanbeige4.2-3B on Agon-Bench: • 84.2% composite • 44.9 tok/s on an AMD RX 5700 XT • 2.4GB VRAM Against the current field: @deep_reinforce Ornith-1.0-9B (5.3GB, 82.7%) Nanbeige wins overall 84.2 → 82.7. Agent is a massacre: 95.8% vs 81.9%. Ornith takes reasoning (98.0 vs 90.2). Code is effectively tied (66.7 vs 68.0). A 3B model matching a 9B on code while decisively outperforming it on agent tasks is not something I expected. @GoogleAI Gemma4-26B-A4B (16.9GB, 84.6%) This is the crazy one. Nanbeige finishes just 0.4 points behind a model nearly 7× larger in VRAM, while running 3.3× faster. Agent is tied at 95.8%. Nanbeige actually wins code (66.7 vs 60.0). Gemma still owns reasoning (98.0 vs 90.2), but the efficiency gap is staggering. @PrismML Bonsai-27B-Q1_0 (3.8GB, 82.2%) Two radically different approaches to “small.” Bonsai compresses a 27B model into 1-bit ternary weights. Nanbeige instead loops its own transformer layers. Nanbeige wins overall (84.2 → 82.2) while running 3.6× faster. The architectural trick is what makes this interesting. Nanbeige has only 22 physical transformer layers, but each layer is executed twice, giving 44 effective passes through the network. Instead of adding parameters, it adds depth. That appears to be enough to produce a 95.8% agent score, competing with models five to seven times its size. The weakness is still knowledge-heavy coding. Regex Engine (33%) and String Cleaner (53%) drag the score down. Layer looping can create more reasoning depth, but it can’t invent training data that isn’t there. This is one of the most impressive efficiency results I’ve seen on consumer hardware. Small models are catching up much faster than I expected. Custom GGUF and @Italianclownz ROCmFPX recipe to run on RDNA at home in replies ⬇️
We released Nanbeige4.2-3B, its Looped Transformer increases model capacity without adding parameters, delivering a capable 3B agent. Nanbeige4.5 is training with LoopSplit, mHC+depth attention & concatenated n-gram embeddings,already in the modeling code. huggingface.co/Nanbeige/Nanb…
1
2
15
1,611
Nanbeige retweeted
🚀 Nanbeige4.2 @nanbeige lands on ModelScope with two models: - 28T-token Looped Transformer base - 3B SFT+RL agent for tools, coding, reasoning, office work & deep research 📊 Team-reported wins over Qwen3.5-9B and Gemma4-12B on agent benchmarks. 🤖 modelscope.ai/collections/na…
14
37
322
20,842
In both LeetCode's Weekly Contests (Weekly Contests 489–491) and the HMMT February 2026 (Harvard-MIT Mathematics Tournament), Nanbeige4.1-3B's performance not only significantly outperformed that of Qwen3.5-4B but also surpassed Qwen3.5-9B.
12
18
206
34,107
The model performance comparison between Nanbeige4.1-3B and Qwen3.5-4B (from its model card).
15
15
235
19,262
Nanbeige retweeted
Replying to @N8Programs
Thank you again for your interest! We hope the model will attract wider attention and be tested by the community to evaluate its performance. The technical report will be released tomorrow—stay tuned! 🌟
3
2
55
10,724
Replying to @nanbeige
Congrats! Now everyone with iOS devices can try the Nanbeige4.1-3B model immediately on their phone. This model excels at tool calling and tends to output many thinking tokens, which requires a large context window. I set 12K context on my iPhone 16 Pro Max with 8K max output, and it didn’t crash. If you have a newer device, you can try higher settings, because this model really thinks a lot. I tested the search_web and fetch_url tools. It’s fully compatible out of the box with existing tool chains. Try it on Privacy AI. It’s completely free for local models. P.S. I ran it on a real iPhone and recorded the video via iPhone mirroring, which is why the refresh rate isn’t great.
1
19
4,236
Nanbeige retweeted
Intriguing new model called 'Nanbeige/Nanbeige4.1-3B' released, appears to be *extremely* SOTA for its size range. So much so that I question if benchmaxxed. But @nanbeige appears to be a small but real lab out of China so I have faith! Quite exciting - will test.
13
24
250
83,353
In the Berkeley Function Calling Leaderboard(gorilla.cs.berkeley.edu/lead…), Nanbeige4-3B-Thinking-2511(huggingface.co/Nanbeige/Nanb…) ranks 25th overall, ranking among the top 10 open-source models and outperforming Qwen3-32B, despite it's only a 3B model.
1
13
2,108
🤖Meet Nanbeige4-3B from Boss Zhipin—a 3B-parameter LLM that outperforms Qwen3-32B on math (AIME), science (GPQA), and tool calling (BFCL-V4), while matching Qwen3-30B-A3B on human preference alignment (Arena-Hard-V2). How? ✅ 23T tokens of ultra-curated data ✅ Fine-grained WSD scheduler ✅ 30M+ high-quality SFT instructions ✅ Multi-stage RL + innovative distillation (DPD) ✅ Chain-of-thought reconstruction & deliberative generation It even ranks top 15 on WritingBench & EQ-Bench3—beating models 100x larger like GLM-4.5 and Deepseek-R1. All models + tech report now open-source: 🔗 Weights: modelscope.cn/organization/n… 📄 Paper: arxiv.org/pdf/2512.06266 Proof that with smart data + smarter training, small can be mighty. 💡
1
26
162
29,604
Replying to @AdinaYakup
Tested Nanbeige4-3B-Thinking(Q3_K_S) locally in Privacy AI with on-device tool calling (search_web). Performance on iOS is excellent. At 3B, it’s lightweight enough to serve as a practical daily offline assistant, yet still handles reasoning and tool use reliably. Congrats to @nanbeige on a very solid release 👏
1
7
1,202