OpenBMB (Open Lab for Big Model Base) aims to build foundation models and systems towards AGI. Connect with us: discord.gg/7q3ry8Ny8K

Pinned Tweet
🚀 Meet MiniCPM5-2B, a 2B-parameter language model bringing high intelligence density to the edge, now open source! It ranks #1 among open-source models under 4B parameters on the @ArtificialAnlys Intelligence Index, with a score of 23. It also scores 20 on the Agentic Index, bringing an early form of general-purpose agent capability to the edge. Across 34 benchmarks, MiniCPM5-2B achieves an average score of 53.9, covering coding, math, long-context understanding, tool use, and agentic tasks. And this release goes beyond the model itself. We’re opening up the data, training recipes, and RL stack behind MiniCPM5-2B. 🤗 Hugging Face: huggingface.co/openbmb/MiniC… 💻 GitHub: github.com/OpenBMB/MiniCPM Modelscope: modelscope.cn/models/OpenBMB… Web: openbmb.cn/
189
318
2,296
655,498
🚀 FIT-GGUF brings controllable-size mixed-precision quantization to MiniCPM5-2B Developer @Scorp1o_117 used FIT-GGUF to build MiniCPM5-2B GGUF variants around specific size and quality targets. Instead of choosing a fixed quantization preset, you can set a target file size or fidelity tier, and FIT-GGUF automatically decides how much precision to allocate to different tensors—then predicts, generates, and verifies the final GGUF. ✨ What’s included 🧠 Tensor-level mixed-precision quantization 📦 Four MiniCPM5-2B builds from ~1.14 GiB to ~1.46 GiB 🎯 Quality / Balanced / Compact / Mini presets 📊 KL Divergence and Same-top evaluation ✅ Generated file sizes matched the predicted targets A nice example of how MiniCPM5-2B can be tuned for different memory and deployment constraints, without being locked into a single Q4/Q5-style quantization preset. Check out FIT-GGUF and try building a MiniCPM5-2B variant that fits your own device budget. 🤗Model: huggingface.co/SC117/MiniCPM… huggingface.co/openbmb/MiniC…
2
1
39
1,499
OpenBMB retweeted
First shot on @OpenBMB based Augury - weeds as indicators app. Time to iron out the bugs 🦠👌 Offline AI, sovereign based, farmer based AI.
2
1
4
350
🎙️ Real-time interpretation, powered locally by VoxCPM2. Developer @HenryZ30734018 built VoxWeft, an open-source simultaneous interpretation system for Apple Silicon. It uses an MLX implementation of VoxCPM2 to turn live speech into translated speech on-device, keeping audio private and responsive. ✨ Highlights: ⚡ VoxCPM2 streams first audio in ~170 ms on an M5 MacBook 🌍 Generates speech across 30 languages, supporting direct language-pair interpretation without a pivot 🗣️ Clones a target voice from ~5 seconds of reference audio for a consistent interpreted voice 💻 Runs in 4-bit quantization on MLX, making low-latency local speech generation practical on Apple Silicon VoxWeft shows how VoxCPM2 can become the speech layer of a full real-time application: not just producing audio, but enabling private, multilingual interaction that stays on-device. Try VoxCPM2 and see what you can build with it! 🔗GitHub:github.com/HenryZ838978/VoxW… 🔗Devlog:github.com/HenryZ838978/VoxW… 🤗 VoxCPM2:huggingface.co/openbmb/VoxCP…
3
2
47
2,004
What can a 2B open-source model actually do on your phone? MiniCPM5-2B from OpenBMB is 2B params, #1 under 4B on the Artificial Analysis Intelligence Index, and tops their Agentic Index outright, 20 vs 9 for the next small model (official). That's the Densing Law in action: model capability density doubles roughly every 3.5 months, and this one fits on an iPhone. The thing I want from a model that size is an iMessage agent: reads my texts, makes reminders and events, looks things up, replies for me, nothing leaves the phone. So I tested it. Six tools, eight real-world scenarios, three runs each. My results, not official ones. Worked: "remind me to send Jordan the invoice tomorrow at 9am" → correct reminder, 6/6. Delta confirmation text → clean JSON of flight, times, seat, 6/6, ~1.5s. "is the 6 train running Saturday?" → searched, read the result, replied correctly: take the 4 or D. "book it and tell her yes" → calendar event plus reply to Maya, 2/3. Didn't: "dinner thursday 7" became Sept 14, not the 10th, in most runs. Once it reasoned the right date and then wrote the wrong one anyway. Decided Lucali is in San Francisco. It's in Brooklyn. Thinking mode ran away: 4,096 tokens on a Swift date parser, no code, three times. HuggingfaceRepo:huggingface.openbmb.cn/model…
19
13
38
157,366
OpenBMB MiniCPM-V4.6 é um dos sinais mais claros de que a IA multimodal está ficando pronta para edge AI. ~1.3B parâmetros. Roda localmente em hardware de consumo. E ainda lida muito bem com OCR, dashboards, gráficos e compreensão de documentos. Testei no iPhone + Mac. 🧵
Paid partnership (ad)
12
19
38
57,525
Thanks so much for sharing this! Really cool to see MiniCPM5-2B being used in a practical multi-agent workflow like this — especially with the workers actually handling matching, short payments, duplicate references, and disputes through tool calls. Appreciate all the work you put into testing and documenting this. Such a nice case for the community 🙌
Next test on Spark 1: 32 synthetic invoices, with purchase orders and payment records. GPT-6 Astra coordinates the MiniCPM5-2B workers. They sort out matches, short payments, duplicate references and price disputes, then write the results into a case ledger. All 32 verified in 67.8 seconds. Eight in each category, with 232 executed tool calls. The video is real time. This demo doesn’t move money.
3
23
1,601
我们期待一起建设的推理框架,能够说清楚优化为什么有效,条件改变后又为什么失效。把一个反常结果追到可以解释、可以复现,再把这些认识写进代码和测试里,本身就是很扎实的系统工作,也正是我希望更多朋友在 SGLang Omni 中共同完成的工程训练。
Article

从一次 Batch Size 争论,思考 SGLang Omni 的性能验证与调度取舍

在《重新审视 CPU 资源作为语音模型 Serving 过程的一等公民》一文中,我们讨论过一个颇为尴尬的问题:同一个 commit

5
4
39
6,161
This is awesome! Running the whole perception stack with a Pi Zero feeding MiniCPM-o4.5 in real time is a seriously cool use case. Love seeing the real-time interactions with the personal robot!
I gave eyes and ears to my robot. I hooked up a Raspberry Pi Zero W2 (webcam + ambient mic), streaming real-time audio/video back and forth to MiniCPM-o 4.5 @OpenBMB open source model running locally on my PC. Zero cloud APIs, full duplex interaction. Project on my git
2
21
1,393
In the AI era, information moves fast. Staying current matters. Developer @markfenner built MiniCPM News Desk, a local-first news briefing system powered by MiniCPM5-2B. It collects official AI and technology updates, identifies the most relevant passages, and turns them into a structured daily recap. Instead of asking the model to summarize everything end to end, MiniCPM5-2B is used to select the key passages from original sources. A deterministic editorial layer then preserves important dates, conditions, and context before the final brief is assembled. ✨ Highlights: 🔎 MiniCPM5-2B extracts key passages directly from source articles 🔗 Every insight stays connected to its original source 🧩 Rule-based checks preserve dates, conditions, and supporting context 🛡️ Invalid or incomplete outputs are rejected instead of silently published 💻 The full pipeline runs locally, with no hosted-model fallback When AI news changes by the hour, the challenge is not just to summarize faster, but to keep up without losing the details that matter. That’s exactly what MiniCPM News Desk is built for: a lightweight, traceable way to turn fast-moving information into reliable daily briefings with MiniCPM5-2B. 🔗 GitHub: github.com/fenner888/minicpm… 🤗MiniCPM5-2B: huggingface.co/openbmb/MiniC…
7
2
38
1,780
OpenBMB retweeted
I wanted to see how useful a small local model could be on my older computer, so I gave @OpenBMB’s MiniCPM5-2B a simple job. The machine: an Intel i5-9400F with 16GB of RAM, running Linux. I built MiniCPM News Desk around something I already spend time doing: keeping up with AI and tech news. ☕ On my machine, it’s connected to an hourly official source collector and configured to send one Telegram recap covering the previous 24 hours. A lead story, other meaningful updates, and quick hits with links to the original sources. Enough information to understand what happened and whether it matters to me. If I want more detail, I can open the article. The code handles collection, filtering, and delivery. MiniCPM selects useful passages from supported articles, while the system preserves important dates, requirements, and limitations. Smaller updates use clearly labeled publisher excerpts or headline links. A small, practical project using hardware I already own. I started with AI and tech because that’s what I follow. The same idea could be adapted to other interests. Here’s what the Telegram output looks like:
6
3
18
3,518
Love this research-agent design from @aijoey: GPT-6 orchestrates while MiniCPM5-2B runs locally on DGX Spark to read sources and extract traceable evidence. A strong example of pairing frontier reasoning with compact, efficient models for transparent research workflows. 🔍
Been putting MiniCPM5-2B to work on my DGX Spark. Built an app called Footnote that turns a research question into a living map of evidence. The setup: • Astra orchestrates through Codex CLI • MiniCPM5-2B workers read sources and extract quoted findings locally • Astra reviews the findings and connects them, with its reasoning visible First question: what’s keeping robots from being useful in ordinary homes? 21 sources. 31 claims. 11 unanswered questions. Click a claim and read the source passage. Click a connection and see why the AI made it. Live tokens/sec and concurrency are visible too. Then we added another paper. The map updated, and the original version stayed intact. All that extraction work from a 2-billion-parameter model. Next experiment: the AEON Qwen 27B model on my second Spark taking over orchestration. Already tested basic tool calls. A full local run is next. Building this one step at a time.
1
3
34
1,681
OpenBMB retweeted
🏥 This is the best reason for a tiny Local AI model. ⚕️ OpenBMB and OpenMed now have a open weight healthcare workflow built around MiniCPM5-2B. 🩺 OpenMed is a whole local-first clinical AI stack: 🔒 PHI / PII detection + de-identification 🧬 Clinical entity extraction 📋 FHIR + HL7 workflows 🔧 MCP + agent tools 📱 Apple / Android / browser support 🌐 33 model-backed PII languages 📚 2,266 model entries After the required models are downloaded, its core processing can stay on hardware you control and even air-gapped. Then MiniCPM5-2B becomes the agent brain. And this is tiny 🧠 2.52B total parameters ⚡ 1.98B non-embedding 📚 131K context Yet OpenBMB specifically trained it for agents: 🤖 500K agent-training samples 🔧 native tool calling 🧠 long-context reasoning 💻 coding + search agents And you can actually run it locally ... ✅ GGUF / llama.cpp ✅ Ollama ✅ LM Studio ✅ MLX 4-bit ✅ GPTQ 4-bit This is why I think small Local AI models may become incredibly important. You don't necessarily need a 70B model reading every sensitive medical record. You may need a 2B local agent sitting next to the data, calling specialized tools and models that each do one thing extremely well. Keep the private data local. Bring the intelligence to the data. 👀 ⚠️ Local processing and de-identification can improve privacy, but OpenMed explicitly says using the SDK alone does not establish HIPAA compliance. 🔗 GH: /maziyarpanahi/openmed 🔗 HF: /openbmb/MiniCPM5-2B
3
3
21
2,001
OpenBMB retweeted
This is my MacBook Air M2 (16GB) with Hermes Agent powered by MiniCPM5-2B. 96k context, and a very respectable output of 20-24 tok/s. This is totally usable, and MiniCPM5 can do all the tool calls to manage the basics with Hermes in a personal agent context. I wouldn't use it to build the next Salesforce, but for the home enthusiast this is a powerful and fast local model that can really punch above its weight, and runs on almost anything. Look at it go.. Amazing!
14
8
137
23,336
Love this run. MiniCPM5-2B on a local Mac, reconciling an old med list against a later note, and correctly keeping naproxen in history instead of the current export — with every graph edge tied back to source. This is exactly the kind of careful, on-device clinical workflow we hoped the model would support. Thanks for putting it to work 🙌 @OpenMed_AI
We gave MiniCPM5-2B from @OpenBMB an old medication list and a newer note saying naproxen was stopped. With OpenMed 2.5, it stays in history, not the current-medication export. Every graph connection links to its source. Local Mac. Fictional notes. Here’s the run 🙂
3
3
17
1,672
MiniCPM5-2B running locally on a 16GB MacBook - beats Qwen3.5-4B on benchmarks and calls web search on its own. A 2B model browsing the web from your laptop. Local AI keeps getting harder to ignore.
atomic.chat
11
7
54
51,654
OpenBMB retweeted
MiniCPM5-2B from OpenBMB is a 2B model with capabilities comparable to a 4B model. This looks perfect as a local model for tool-calling that can saves you cloud tokens.
4
2
17
1,958
MiniCPM-o 4.5 meets real-time voice agents. Developer @mrgoodmantweets built MiniCPM-o Booking Desk, an open-source appointment booking agent that puts MiniCPM-o 4.5 at the center of a real-time, full-duplex interaction workflow. With MiniCPM-o 4.5 handling continuous audio-visual interaction, the agent can listen, speak, and read live booking status from an operator screen while completing tasks in real time. ✨ Highlights: 🎙️ MiniCPM-o 4.5 enables natural, full-duplex voice interaction 👁️ Multimodal perception connects the conversation with the live booking state 🧠 Model-driven interaction is paired with deterministic state control for reliable execution ✅ User confirmation is required before an appointment is actually booked A practical example of MiniCPM-o 4.5 powering agents that can perceive, interact, and take action in real time. 🙌 🔗 GitHub: github.com/AlessandroBonomo2… 🌐 Project write-up: alessandrobonomo28.github.io… 🤗 MiniCPM-o 4.5: huggingface.co/openbmb/MiniC…
Made with AI
4
15
75
3,096
This is the kind of local AI use case I love seeing. Not just “AI can identify a plant,” but actually helping farmers make better decisions from a photo — entirely on-device. 80.2% → 90%+ top-1 is a big jump, but getting this into a simple phone GUI for farmers is where it gets really interesting 🌱
Local model - Augury, that genuinely help farmers make better farming decisions: Some evening tweaks on Augury @OpenBMB model. Here are the highlights, I believe we will reach 90%+ plant id accuracy by the end of the weekend. - Photo ID went from 71.8% to 80.2% top-1 by merging duplicate species keys and adding PCA whitening. - The formatter invented pH numbers in 7 of 20 answers; it now invents none because a structural pass catches and neutralises them. - The old 0.80 confidence threshold was backwards — precision peaks at 0.65 and falls above 0.75. - Australian species with soil data went from 188 to 221, including Paterson's curse and serrated tussock. - The extractor now resolves all 11 names that previously failed, via 158 curated common names. - Management questions like "what herbicide for lantana?" now decline instead of answering off-topic. - Two claims in my earlier report were wrong and are now corrected with visible notices. - 90% top-1 was not reached, and the measured reason is the encoder on CPU, not the method. We keep on going! Weeds as indicators model is improving. Next step - create a GUI wrapper so it can be shipped to those who will use it - farmers on their phone.
5
1
12
1,271
This is a really practical way to use a small local model. Also pretty impressive that the whole pipeline runs CPU-only through OpenClaw + Ollama. It’s a good example of how workflow design can make a small model much more useful than the raw model size would suggest.
From Raw Logs to Root-Cause Clues What happens when a small local AI model gets a real debugging job? #OpenClaw #Ollama #AI #LLM #DevTools
3
2
18
1,171
This is exactly the kind of small-model use case we love. No massive GPU, no cloud API, just a 2B model turning an old PC into a useful personal news desk. Small models don't need to do everything. They just need to do one useful thing reliably. ☕
I wanted to see how useful a small local model could be on my older computer, so I gave @OpenBMB’s MiniCPM5-2B a simple job. The machine: an Intel i5-9400F with 16GB of RAM, running Linux. I built MiniCPM News Desk around something I already spend time doing: keeping up with AI and tech news. ☕ On my machine, it’s connected to an hourly official source collector and configured to send one Telegram recap covering the previous 24 hours. A lead story, other meaningful updates, and quick hits with links to the original sources. Enough information to understand what happened and whether it matters to me. If I want more detail, I can open the article. The code handles collection, filtering, and delivery. MiniCPM selects useful passages from supported articles, while the system preserves important dates, requirements, and limitations. Smaller updates use clearly labeled publisher excerpts or headline links. A small, practical project using hardware I already own. I started with AI and tech because that’s what I follow. The same idea could be adapted to other interests. Here’s what the Telegram output looks like:
2
2
43
2,293