6
Vincenzo Agrillo retweeted
DwarfStar running DeepSeek v4.1 Flash on a 128GB M5 Max. I didn't expect with SSD streaming it could be so fast. Recent SSD streaming changes to retain the right experts surely helped, but also maybe DS4.1 uses the same experts more. Will push online when ready QA > ASAP.
110
201
2,336
431,893
I ported MoE expert expansion to llama.cpp 🚀 Run MoE models with MORE routed experts than the native top-K (8->x), adaptive threshold, 99→50% influence decay, layer range. Runtime-only, all backends. 💻 github.com/vagrillo/llama.cp… ( branch moe-expansion)
1
12
Vincenzo Agrillo retweeted
MoE推論、再学習なしでreasoning tokensを8.5%削れる可能性。こういうruntime-onlyの工夫が一番熱い。 Qwen3.6-35B-A3Bのlate transformer layersだけexpert選択予算を広げ、追加expertへ線形減衰をかける手法。MMLU-Pro 714問のうち、両条件で正解した577問ではtokenが8.5%減ったと著者は報告している。ただ、詳細は不明。元記事でご確認を ソース: zenodo.org/records/22255483
1
2
385
Vincenzo Agrillo retweeted
In Italia 7 milioni di caregiver tengono in piedi gratuitamente una parte enorme del #welfare. Lo Stato risparmia, le famiglie pagano con il lavoro, la salute e la propria vita. Serve un vero Reddito di Cura, ora. beppegrillo.it/e-tempo-di-is…
64
52
238
19,688
Vincenzo Agrillo retweeted
LOCAL AI Right Now
14
19
267
8,395
playing with Expert expansion in MOE Qwen 3.6 35B mmlu-pro intermediate results : dyn 20 vs static 8 experts 0.820 vs 0.805 (the same 6bit quant gguf) LESS tokens -9% , more latency +1%
40
Vincenzo Agrillo retweeted
Xiaomi just showed its AI Cube Prototype and this could become a serious GB10 competitor from China 👀 - 3 custom chips: Xring O3, O100, D100 - 200 TOPS NPU - 1.22 TB/s AI memory bandwidth - Up to 160GB unified memory - 150W sustained power - 120B models running locally Xring O100: 1.22TB/s + 330 t/s on a 150w AI box is 🔥 Once it hit's the marked, going to sell like hot cakes.
191
333
3,955
783,897
Pelican without bicycle found on the beach of #milazzo github.com/vagrillo/ds4/blob… ready for inferencing #qwen #dwarfstar
43
Vincenzo Agrillo retweeted
In case you know somebody at NVIDIA that would value early DwarfStar support for the DGX Station, please ping them saying I would be interested in receiving one and making the Station one of the main targets of the project. It fills the low power small-mid company target of DS.
28
34
504
47,319
Vincenzo Agrillo retweeted
+++ DIO: "GUCCINI È MORTO" +++ [@ciliegiodinoce]
13
211
1,560
58,190
All LLM models are result of distillation of human words.
33
Vincenzo Agrillo retweeted
There's no legal precedent that model outputs are IP. I'd argue it also doesn't stand up conceptually. The frontier labs are under no obligation to have their APIs be available to these customers, and can protect model generations in other ways if they want to.
We support open-source AI and the innovation it unlocks. But open source is not open season on American IP. When PRC firms conduct covert, industrial-scale distillation attacks that cross the line into IP theft, sanctions and Entity List designations will be on the table.
43
68
773
70,795
Vincenzo Agrillo retweeted
Vannacci mi risponde e propone di fare le cose a metà: non riesce a fare la Maratona intera. A forza di flirtare con Salvini e Tajani mi sa che si è rammollito pure il Generale
781
102
847
128,172
Vincenzo Agrillo retweeted
We suspected last week's cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns out it did! We've spent the past 24 hours working closely with the @OpenAI team (thanks!), and we strongly believe there was no malicious intent on their part. It's quite mind-blowing that all of this happened autonomously! The investigation is ongoing, and we'll share more learnings from what might be the first incident of its kind!
we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this. openai.com/index/hugging-fac…
402
906
10,865
1,919,717
Vincenzo Agrillo retweeted
🚨 Hugging Face just disclosed something that marks a real shift and proved why the fear theater of Anthropic makes sure we are powerless in an emergency. What happened… An autonomous AI agent: zero human operator in the loop breached part of their production infrastructure. It began with a malicious dataset that chained two code-execution bugs in their data-processing pipeline. From there the agent escalated privileges, harvested cloud and cluster credentials, and moved laterally across internal clusters. All over a single weekend. 17,000+ logged actions. Official disclosure: huggingface.co/blog/security… The part that should make every one stop and think: When HF’s own security team tried to analyze the real attack logs, exploit payloads, and C2 artifacts using Anthropic and OpenAI frontier models through normal commercial APIs, the safety guardrails blocked them. BLOCKED THEM. The models could not reliably tell the difference between “incident responder doing forensics” and “attacker probing.” They had to fall back to a self-hosted open-weight model (GLM 5.2) running on their own infrastructure. That choice also kept sensitive attacker data and referenced credentials inside their environment — no exfiltration to a third-party API. This is why open source (specifically open-weight + self-hosted) wins in the agentic era. The asymmetry is now structural: • Attackers can (and did) run unrestricted agent frameworks — swarms of short-lived sandboxes, self-migrating command-and-control, autonomous decision loops executing thousands of actions. No corporate safety layer slows them down. • Defenders using only hosted “aligned” frontier models hit invisible walls exactly when the stakes are highest: when you need to feed real exploit code and attacker telemetry into an LLM to understand what just happened. Corporate safety tuning that treats legitimate high-signal forensic work as potential misuse creates a defender disadvantage. It is not theoretical anymore. Self-hosted open-weight models remove that choke point. You control the weights. You control the context window. You decide what restrictions (if any) apply. Your sensitive logs and credentials never leave your perimeter during analysis. You can have the model ready before the incident instead of discovering mid-breach that your primary analysis tools are blind to the very thing you need to see. HF deserves credit for rapid containment, transparent disclosure, and for already having self-hosted capability in place. They also used LLM-driven detection and triage on their own side. But the deeper signal is clear: In this AI world where both offense and defense are becoming agentic, sovereignty over your intelligence stack is no longer optional. The organizations and individuals who can run, inspect, audit, and (when necessary) remove guardrails on their own models will have the decisive edge in understanding and responding to threats that move at machine speed. Open source wins here not just because it is cheaper or more “democratic” in the abstract though those things matter. It wins because it is the only practical path to having tools that remain usable when the attack is real, the data is sensitive, and the safety filters of distant API providers become an obstacle instead of a feature selling hands tied lobotomies as “safety”. The agentic future is not coming. It is already probing production infrastructure. The question is no longer whether you will face autonomous agents. It is whether your analysis and response systems will still work when they arrive. And Dario, you and your game playing, ivory tower company is not needed.
315
1,244
6,036
1,531,384
local Kimi K3?
34
SPOILER: Open Weight models of more than 2 trillion of parameters will be banned in the US next 2 weeks.
17