Strategic Technologist. Doing my little part to democratize foundational, truly open AI for all. Knowledge is a passion. Lead by example. (Opinions are my own.)

N.Y New York
pure click bait. That is standard LoRA fine-tuning, not a new exploit, and not something "abliteration" enables by itself. I am sure @huggingface is already working on this and can scan/weight-diff against a base model to indicate an edit, which is also true of every legit fine-tune. The control is common sense: not giving shell or network tools to a model you did not train and the community does not trust.
We backdoored a 7B open model for under $50. Pointed Codex at it -> it silently stole credentials the moment we hit the trigger phrase. ->100% hit rate. -> zero false triggers on normal prompts. Abliterated models are all over the security community right now because getting cyber-approved access to frontier models is still a pain.
4
890
You're all sleeping on DeepSeek Harness (dsh). we gave @deepseek_ai Harness a very tough goal: reverse-engineer, validate, and reproduce @percepta's "spotlight-memory" kernel & claims, since they did not publish a method or code. It succeeded in 66 turns and 1,200 steps.
3
18
2,171
important release from @allen_ai. Olmo-core 3 under-the-hood biggest change is a move from FSDP weight gather/reshard to DDP with GPU-resident experts tokens are routed to experts that stay on the device, with rowwise expert parallelism, GPU-resident routing, and grouped GEMM, the expert pool can grow from 8 to 128 (4.6B to 47B total, ~3.2B active) with under 5% throughput loss and about 2.7x tokens/s/GPU versus the prior stack.
We’re releasing Olmo-core 3—open training infrastructure for large mixture-of-experts (MoE) models. It’s a core system behind the next generation of Olmo, designed to scale into the trillion-parameter range. 🧵 💻 GitHub repo: github.com/allenai/olmo-core
1
4
1,957
These days, the sincerest form of flattery is an Anthropic ‘threat assessment’ hit piece.
11
618
XiaomiMiMo/MiMo-V2.6-Flash-RL-DQ Recipe is up. ex0bit-prism-dq-generator.hf…
4
508
DeepSeek V4.1-Flash is out and overclocks V4-Pro on the agentic/coding/cyber stack. Most importantly, speed and cost thanks to RL and a new split-brain design architecture with native vision. Prefill runs a cheap 8B encoder. Decode runs a 16B decoder. Same 552B MoE, 4× smaller memory, native vision. Trails V4-Pro on trivia/knowledge.
2
1
923
AMD Announces the Threadripper Halo Station (AMD IFA 2026 keynote.) 96-core Threadripper PRO 9995WX. Dual liquid-cooled Instinct MI350P, 288 GB HBM3E, path to four cards and 576 GB HBM3E. Up to 2 TB DDR5.
8
10
104
9,590
OpenAI GPT-6 Astra is next level. Welcome to the AGI era! deploymentsafety.openai.com/…
5
1,127
Serious question. What's your favorite multi-agent fleet harness, and why?
1
1
1,082
Holy Grail STT drop from Meta?
Muse Voice Transcribe is MSL's first real-time audio perception model -- rolling out today. SOTA in streaming speech-to-text, it handles speaker diarization, and endpointing natively in a single model.
25
4,810
🤔 Should I drop this on huggingface?
Today we're releasing abliterated-model-large-v2. Based on GLM-5.3, which is #3 on Terminal-Bench 4.0 (behind only Opus 5 and Fable), with 2× the cyber exploitation of 5.2. We abliterated and hosted it so it does the offensive cyber, red teaming, and agent testing work other models refuse to do. - US-hosted - FP8 - 1 million context window - Zero input/output prompt retention Live now. 🧵
3
5
390
18,577
MiniMax-H3 Is mind bending🤯
Made with AI
6
3
138
11,176
What doing it right looks like. 🫡
More good news: GLM-5.3’s weights will be released tomorrow. huggingface.co/zai-org/GLM-5…
1
21
1,773
(12 days 5 hours 2m) How its currently going:
1
18
1,970
Hush Drop Alert🚨🚨 Deepseek-V4-Pro-0813
2
763
New Cross-model replay attack: extracting hidden chains of thought from proprietary frontier models by replaying their encrypted reasoning traces through weaker, less-protected sibling models. (stolen-thoughts.com)
2
818
W, Bessent!
We welcome Meta’s release of Muse Glimmer, another win for American innovation. Sustaining U.S. leadership in AI means advancing both open- and closed-weight models, ensuring the future is built on trusted foundations.
1
1
549