1/ Agentic LLMs can automate vuln detection. Very exciting, but doesn't address the hardest part (imo) of vuln research: prioritization. Can we reliably explore the search space and separate signal from noise? I wrote a paper (and OSS tool) to solve this. arxiv.org/pdf/2512.06155
2
60
217
106,630
2026, at the height of medical progress: the vibe pediatrician prints out this anatomically questionable image and proudly thinks, "welp, LGTM!" 🙃 can't wait for 2027.
1
5
574
Caleb Gross retweeted
I'm hiring an exceptional Offensive Security Researcher for my team at NVIDIA (Offensive Security Research - OSR). Firmware, microcode, RISC-V, hypervisors, and shipping mitigations like HW CFI, Memory Tagging, and Pointer Masking from the ground up. jobs.nvidia.com/careers?quer…
13
83
469
44,825
cyber fine-tune *must* be called Sev
UPDATE: Kev-0.6B, 4B, and 8B are now available. Kev is a family of small open source Jev-like decision models you can train and run yourself. This new family is based on Qwen3 using the same LoRA + small pointer head technique as before, but scaled up. Out of domain, on data Kev never trained on: Kev-8B 79.6%, Jev 85.7%. • Drop-in TypeSafe System One API; their SDK works with one `base_url` change • Kev-4B serves on a 32 GB Mac in bf16: ~300 ms for five questions, ~40 ms on an H100 • Repeated documents hit a KV cache: 2-2.5x faster • Apache 2.0 License. Kev-4B trains in 40 minutes on one H100. Kev-8B in 83 minutes. Code, weights, evals: github.com/jaredpalmer/kev
1
6
755
Caleb Gross retweeted
The funniest part about local AI: you start because you want privacy A few weeks later you're comparing memory bandwidth, quantization formats and $5,000-$20,000 possible hardware purchases at 2am
134
117
2,224
100,497
Caleb Gross retweeted
Jev launched Monday as a closed API. By Thursday we already have an open weight alternative we can run locally. This is fun, looking forward to test it! huggingface.co/convaiinnovat…
19
42
385
22,937
Caleb Gross retweeted
Absolutely incredible work by @mdowd and @UnpromptedAU putting together an amazing collection of people and talks. Favorites were @chompie1337 and @justdionysus – great speakers with deep expertise in exploit dev. Hope to come back next year for another round!
1
5
57
2,863
Caleb Gross retweeted
Umbriel's Caleb Gross (@noperator) has an article in the latest Phrack 73 issue (@phrack) "Word Machines for Weird Machines" - check it out when you get the chance!
5
20
4,056
been driving Hermes today on Qwen3.8-Flash-Next via pwilkin.github.io/strix-halo. just amazing.
Replying to @antirez
I've been running DeepSeek V4 Flash 0731 through DS4 engine on the Framework. With a Q2-class quant, I was getting roughly ~200 tok/s prefill and ~15 tok/s decode. I did not feel that this configuration was fast enough to work well in an ongoing Hermes loop. Qwen3.8-Flash-Next is a much smaller model while still roughly on the same capability level of DeepSeek V4 Flash 0731 (even better on some tasks), and I'm running it at a Q4-class quant. Wilkin has been aggressively optimizing llama.cpp/ROCm specifically for Qwen3.8-Flash-Next on Strix Halo. On my Framework I'm seeing roughly ~1,300 tok/s prefill and ~30 tok/s sustained decode at around 131K context.
5
1,435
Earlier this summer, I bought a 128 GB Framework Desktop just so I would have access to a large continuous block of unified memory to run frontier-adjacent models at home. I've been really inspired by @antirez's DS4 project which aggressively targets specific model+hardware combos. I recently became aware of @ilintar's project doing this specifically for Strix Halo and I'm really excited about it.
3
3
27
4,406
I've been running DeepSeek V4 Flash 0731 through DS4 engine on the Framework. With a Q2-class quant, I was getting roughly ~200 tok/s prefill and ~15 tok/s decode. I did not feel that this configuration was fast enough to work well in an ongoing Hermes loop. Qwen3.8-Flash-Next is a much smaller model while still roughly on the same capability level of DeepSeek V4 Flash 0731 (even better on some tasks), and I'm running it at a Q4-class quant. Wilkin has been aggressively optimizing llama.cpp/ROCm specifically for Qwen3.8-Flash-Next on Strix Halo. On my Framework I'm seeing roughly ~1,300 tok/s prefill and ~30 tok/s sustained decode at around 131K context.
3
4
2,039
Check out Piotr's work here: pwilkin.github.io/strix-halo
1
316
Caleb Gross retweeted
an image encoder/decoder library heap overflow bug landing you RCE on a trillion dollar company is a 2013-era, AFL-style fuzzing fantasy
1
7
147
5,851
Caleb Gross retweeted
You can trust it with your life.
1
6
1,052
Excited to announce that I recently joined @umbriel_ai with @mdowd and @dyn___ :) Humbled and grateful to get to work with both of these talented hackers.
Exciting announcement: I am launching a new AI-focused company today named Umbriel (@Umbriel_AI)! Excited to do some cool research in the space 😎
9
3
58
5,236
I am drawn to people like Mark and Aaron who, rather than displaying cognitive surrender in the face of highly capable AI, still yield to the Hands-On Imperative in pursuit of ever-deeper understanding.
1
7
262
We'll be working together to build technologies that materially advance what is possible in cybersecurity.
1
4
156