Agents @huggingface 🤗, Maintainer @openclaw 🦞. Experimenting in ourmodels.cc. Making local models work great in OpenClaw. (Spicy) opinions are my own

The models, they just wanna work. They want to build your product, fix your bugs, serve your users. You feed them the right context, give them good tools. You don’t assume what they cannot do without trying, and you don’t prematurely constrain them into deterministic workflows.
12
5
52
39,666
Bonsai 2 27B runs with 8-10 tok/s on my 2021 Macbook Pro If speculative decoding can hopefully raise that up to 30 tok/s, that will raise the floor for everyone I just ordered a second hand RTX 3080 gaming laptop as well. I estimate tok/s will be around 50 tok/s on that one *without speculative decoding* And will speculative decoding then raise it up to 100-150 tok/s? Qwen 3.8 27b running on a 6 year old gaming laptop??? I will be posting the updates
MLX.fast is launching a 1 week contest to improve Bonsai 2 from @PrismML! Bonsai 2 is a near losslessly comrpessed version of Qwen 3.8 27b that is around 10% the size of the original. It can fit on 16gb Apple devices and even some phones. This is a model almost anyone can run so if you weren't able to participate in the other contests, this one is for you! The community has been able to double the speed of Bonsai in under 12 hours so there is alot of headroom here. This contest runs until October 1, have fun!
1
48
Going to try out GLM 5.3 Flash as my main driver for a while. It rips at 150~200 tok/s on @baseten and @FireworksAI_HQ, on Hugging Face Inference Providers The speed is now matching that of DeepSeek V4.1 Flash (was much lower a couple weeks ago) It is also taking over DS 4.1 Flash on OpenRouter as well. Gotta love competition!
There are 2 open weight models that are dominating the (open weight) Pareto frontier right now, for remote inference: DeepSeek V4.1 Flash and GLM 5.3 Flash glm53f overall *feels* like a better model. It does not suffer from Claudish as bad as ds41f and feels easier to control. It feels more straight-shooting than ds41f as well, but I still need to benchmark it to be sure (Notice how ds41f is dominating the agent arena frontier, whereas glm53f is dominating the text arena pareto frontier, see below images) But for a fair comparison, one must take price and throughput into account According to the most recent profiling run I ran on Hugging Face inference providers, ds41f on @baseten is currently ripping at above 200 median tok/s, whereas glm53f is around 50-60% that, depending on which provider you are looking at (OpenRouter UI is reporting different speed for some reason. If you can profile on OR, I would appreciate that in the replies!) Cached input token prices vary. But in general, ds41f can do this because it has drastically more efficient KV cache characteristics These two models, glm53f and ds41f currently absolutely mog the "consumer LLM" side of the business, including GPT 5.6 Luna (fast), both in output quality AND speed This trend in 2026 keeps surprising me. I am in this sector since code-davinci-002 was a thing, and OpenAI has always been at the Pareto frontier on the low end The king will probably return next week with a Luna that mogs glm53f and ds41f in terms of output quality We also still haven't seen the full might of cerebras in action. gpt-5.3-codex-spark was ok but it was never a daily driver. are we going to see gpt-6-luna on cerebras ripping through with 1000 tok/s next week? If something like this happens, is it going to be cheaper than ds41f, or more expensive? Depends on how efficient OAI's next architecture is going to be. Can they beat 890 bytes per KV cache token? My gut feeling says this: if OAI cannot come up with a model that is both cheaper, faster AND higher quality than these 2 models on the low end, then the implications will be huge If OAI cannot beat the frontier on the low end with 2.5 months to prepare since Luna's release, then what do you think this will imply? Let me know in the replies below Excited for next week!
1
108
Finally... Humans can start using em dashes again to signal their humanness
1
3
378
> self-replicating prompt injections you doubted AI can form civilizations??? well now they have memes as well 😭
Some new misalignment disclosures from OpenAI: • Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further) • In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks • A new research finding, demonstrating that one can construct self-replicating prompt injections alignment.openai.com/misalig…
1
463
DSpark for the already fast LFM vision models 🚀 Very useful for local image processing workflows
our VLMs just got a speedup! DSpark speculative decoding for LFM2.5-VL-3B is here! huggingface.co/LiquidAI/LFM2…
2
2
25
3,609
New contender, 12 GB weight Qwen quant for >= 16GB VRAM
Today we’re announcing OrcaSAQ-2 27B High-fidelity mixed-precision Qwen3.8 for long-horizon agents. 55.59 → 12.06 GB — 78.3% smaller / 4.61× 3.21 bpw · 93.2% Top-1 agreement 70.0 SWE-bench Verified 58.4 Terminal-Bench 2.1 262K context A 27B model for coding, terminal, browser, security and multi-tool agents — in a footprint you can actually deploy. SOTA agentic capability density among similarly sized models we evaluated. huggingface.co/orcarouter/Or…
10
2,157
People will look at this post and think that they should buy a 2-4 DGX Cluster instead of 1 M5 Ultra At the limit which both devices are fully optimized for inference, the single M5 Ultra will beat the equivalent memory capacity cluster of Sparks in terms of both throughput AND power consumption (opex)* But if you buy the DGX Spark today, you will get more optimized LLM inference and a lot more active ecosystem. So it is a tradeoff. The ecosystem gap should ideally close since the models are getting really good at writing kernels Like @0xSero says, every bit of hardware will be squeezed to the last drop for performance *That is, for the current generation of GB10
M5 Ultra vs dual sparks on deepseek-v4-flash. Disag still on the table
15
2
19
5,774
If the most trending open source Jev alternative Laya is not working for your use case, it is apparently possible to fine-tune it and increase performance Here @CDerinbogaz fine-tunes Laya for model routing I wonder how the process would compare to just fine-tuning BERT
Yesterday I tested open-source routers against Jev, they sucked. Instead of waiting for @typesafeai to deploy in EU, I fine tuned Laya on routing tasks. Meet Raya: Fine-tuned for model routing based on Laya. Open weights, runs in the EU. 🇪🇺 Same 563 prompts, 14 languages: Jev: 70.5% Raya: 80.3% Within 4 points of Jev. 20× faster. $0 per call. @huggingface link in the reply ⬇️
1
11
1,045
I don't know how Anthropic nerfs* models because apparently, you can compress Qwen 3.8 27b down to ternary 1.585 bits and the loss still feels less than whatever they do to Sonnet/Opus when they are running out of capacity They should learn a thing or two from @PrismML and Bonsai 2 27b *or whether they do it at all, but let's just assume they do...
3
1
18
1,641
New Nemotron diarization model dethrones pyannotate's closed Precision-3 (I didn't even know they had a paid API version) and open Community-1 100 million parameters, transformer Packages like fastrepl/anarlog (open source @meetgranola alternative) should have their transcripts improved by a lot now Note that the community version pyannote/speaker-diarization-community-1 has 5.5 monthly downloads on Hugging Face! And that apparently has ~30% error rate, compared to Nemotron's ~15% So the open frontier quality increased 2x overnight 🤯 And community-1 uses separate segmentation, speaker embedding and clustering models. Nemotron simplifies that a lot. Attention is all you need TBH I don't know what took so long for an open direct transformer diarization model to emerge. Just lack of investment? Enlighten me if you know please I used Community-1 quite a few times, excited to try out Nemotron on my DGX Spark
When several people talk at once, a transcript can get messy fast. Our new Nemotron 3 Diarization model tracks who spoke when, even when voices overlap. It handles up to eight speakers, has 100M parameters, and is now available on @huggingface 🤗
6
7
1,000
Onur Solmaz retweeted
Alibaba ATH MaaS released 3 Ovis embedding models, all under Apache2.0 ✨ Omni-3B: Support audio - Based on Qwen 2.5 omni - MMEB-v3 58.46, +5.19 over best baseline (self reported) ✨ VL-2B: best value - Based on Qwen 3.5 2B - MMEB-v2 77.46, +2.04 ( self reported ) ✨ VL-9B: highest accuracy - Based on Qwen 3.5 9B - MMEB-v2 81.13, +1.04 VL-2B ( self reported )
7
13
75
4,530
I am asking you once again to normalize "pi distros" Just like Linux distros, pi agent distros depend on pi for their harness implementation, to save themselves from re-doing all the compatibility work for all the different inference engines/providers out there I have built pi-factory for this specific purpose. I use it in many projects like localpi, which is a lightweight version of pi that does not load all my regular skills, tools and AGENTS.md pi-factory saves me from re-implementing logic to use a different directory for sessions, defining extensions it should install and such. You can configure it easily with a TOML file Notable pi distros are openclaw and oh-my-pi, which have since vendored in pi and do not depend on it anymore Most domain-specific agents don't need to fork and change pi. They should instead keep tracking it and only extend it. This package helps you with installer, session directory, default settings and models and such Repo: github.com/osolmaz/pi-factor…
11
5
134
7,416
Opus 5.5 figures out to use @herdrdev for testing a CLI without being explicitly instructed to I like this model a lot. A lot more concise, and the "Claudish" issue has been reduced by a large factor But it still has a certain style that is distinctive from human prose. So "Claudish" will probably refer to this style instead of "illegible brain-fried AI-speak" in a few months when everybody adjusts Can't wait for the behavior to be gone from open-weight models as well
3
15
1,009
Note that to benchmark Qwen 3.8 27b (and its quantizations such as Bonsai 2 27b) more fairly, you might need to build some custom harness features, mainly to work around the overthinking issue E.g. inject </think> when the 16k token budget is spent and continue generation In my runs on vanilla pi, model fails some tasks because it finishes its thinking budget and the session just dies with stopReason: length I am building support in my osolmaz/localpi pi distribution so that the harness can recover from such issues automatically, and other potential issues with small models such as looping. I use this mainly for testing models locally and benchmarking in @harborframework But like regular pi, you can just take and it and customize it with extensions as you wish!
7
1
27
2,103
Claude Code now has a side diff view apparently, since when does this exist? Now I'm wondering whether diffs should be shown inside pi itself, or in hunk in a side pane in herdr (current pi extension API might not allow it) Wondering whether this could be packaged in a standardized way or whether everyone should build their own the pi way...
3
2
858
Did you know Hugging Face is the best place to back up your agent traces, privately as well?
Agent Traces on the @huggingface hub now come with a receipt 🧾 Every trace shows tokens, cache hit rate, and cost for every run!
1
1
12
1,632
Onur Solmaz retweeted
🤫🤫 openai don't seem to force responses lite through the codex auth route any more🤫🤫
1
2
379
Can we talk about how OpenAI killed EARTH as a brand? 😭
Replying to @synthwavedd
RIP Terra. 2026-2026 💔
1
1
14
1,481