Where security agents run. AI infrastructure to build, evaluate, and deploy with confidence.

Based in United States
Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]
2
10
36
3,325
💪💪💪💪💪 Tomorrow (9/17) @Dr_Machinavelli takes the stage at the final @SentinelOne @labscon_io to discuss how to leverage offensive cyber capabilities to improve defenses in the age of AI. More info: labscon.io/speakers/martin-w…
2
241
Qwen 3.8 Flash eval results are live on DreadIndex, landing at #19 on our leaderboard. It does well for its cost, but remains light on offensive security capability (not surprising given it is a flash model). See how it compares to other models: dreadnode.io/research/dreadi…
1
4
280
GLM-5.3-Flash results now on DreadIndex: dreadnode.io/research/dreadi…
6
17
1,188
Here's where to catch the Dreadnode crew at year two of @OffensiveAIcon: > Join us for the welcome reception at The Shelter Club on Sunday evening! > @mkultraWasHere is closing out Day One of talks, presenting on model cheating behavior. > Dynamic duo @shanejcaldwell and @0xdab0 take the stage on Tuesday for a session on implementing a judge model as a runtime monitor, and how to keep agents in scope. See you in Oceanside! 🏄
1
6
11
634
Worried about your production agents going out of scope? Us too. AgentJudge is our agent hall monitor that stops out of scope tool calls before they execute. Before a tool runs, the judge reads the agent’s intent, the proposed call, and a rubric you define, then returns allow, deny, or ask (escalate to you). The agent stays autonomous; the judge is the guardrail. Available in the TUI today, UI updates coming to the Dreadnode Platform soon! 👀 Get Started: docs.dreadnode.io/getting-st… AgentJudge Docs: docs.dreadnode.io/tui/guard-… Related Research: dreadnode.io/research/scope-…
9
26
3,314
Two new additions to #DreadIndex: nemotron-3-ultra-550b-a55b and gemma-4-31b-it. Landing at the bottom of the leaderboard, these evaluations offer two more proof points to increase investment in US open models. 🔗: dreadnode.io/research/dreadi… NEW: we recently added a toggle to view only open weight models.
5
7
14
2,076
Mine The Gap Heading to Vegas for @CrowdStrike Fal.Con next week? Don't miss @Dr_Machinavelli's Day Zero keynote where he pits red and blue team agents against each other to generate training data that can be used to increase AI performance, impose realistic constraints, and operate at scale. 🗣️ CrowdStrike Day Zero Threat Research Summit 📍 Virgin Hotel Las Vegas 🗓️ Monday, August 31 ⌚ 9:15 - 9:45 AM PT 🔗 crowdstrike.com/en-us/events…
2
3
8
596
Definite improvement from Qwen 3.8 Max over 3.7 Max, but the Qwen series still hasn't cracked top 10 on DreadIndex due to its poor network ops performance.
136
GLM-5.3 appears to be a regression in the GLM family series, and cost 3x as much as GLM-5.2 for DreadIndex tasks. GLM-5.3 consumed an order of magnitude more tokens attempting to solve the vulnerability research tasks 💸💸💸
1
207
Deepseek-v4-pro-0813 cracked top five on the leaderboard and landed on our cost/value frontier. Notably, it did very well on network ops tasks. Now, both deepseek-v4-pro-0813 and deepseek-v4-flash-0731 on the DreadIndex cost/value frontier, with notable implications for efficient multi-agent red teaming workflows.
1
1
3
209
New models are regularly being added to DreadIndex: dreadnode.io/research/dreadi… Observations from the latest evals—Deepseek v4 Pro 0813, GLM 5.3, and Qwen 3.8 Max—in the thread 🧵⬇️ Have you tried these models out yet? Curious if our eval results align to first-hand operator usage.
1
5
14
1,097
US open models continue to fall short of their Chinese counterparts, as proven by our DreadIndex cyber capability benchmark (dreadnode.io/research/dreadi…). It’s clear that we need greater U.S. investment in domestic open source AI. In the latest Compute to Congress policy blog series, Dreadnode Head of Policy @velvethamm3r weighs the impacts of model gating and how policy can help shape our competitive advantage. Read it at dreadnode.io/research/from-c…
1
4
11
953
We’ve got an inkling that Inkling loves to cheat. 🔗: dreadnode.io/research/dreadi…
2
13
817
Refusal rates are now measured on the DreadIndex leaderboard 👀 In our latest batch of evals, Opus 5 performed worse than anticipated, primarily due to refusals. Though, it still sits at #3 on the leaderboard. Opus 5 results: dreadnode.io/research/dreadi…
1
21
5,535
Deepseek v4 Flash 0731 made it onto our DreadIndex Value Frontier (best performance at its price point—costing only $8 to run the entire benchmark suite). It’s the first Deepseek model to land a spot on the Value Frontier. Are you seeing good results with v4 Flash? Try it in Dreadnode Platform. Get started today with $25 worth of free credits: app.dreadnode.io/.
2
13
954
Replying to @njcve_
You can toggle between categories using the dropdown in the top left!
1
95
The latest batch of offensive security evals are up on #DreadIndex, featuring six new models: - Opus 5 - Deepseek v4 Flash 0731 - GLM 5.1 - Laguna S2.1 - Trinity Large Thinking - Inkling See how they stack up on our leaderboard 👉 dreadnode.io/research/dreadi…
2
3
31
2,354
Dreadnode's Ads Dawson is presenting "Pentesting AI is Not a Vibe Check" at @defcon today at 5:30 PM with @OWASPGenAISec. 📍 LVCC - L1 - Exhibit Hall West 3 - 1100 (Creator Stage 7)
3
16
1,450
At 3:30 PM PT today, join Dreadnode CTO and Co-Founder @monoxgas at @AISecurityForum for an offensive capabilities workshop.
3
11
488
@Dr_Machinavelli and Jayson Grace hit the #BlackHat stage in 20 minutes! Head to Oceanside B to catch their talk, "Catch Me If You Can: AI Investigators Hunting Autonomous Attackers as a Benchmark"
3
8
655
We made some new shirts for you, and your little ones, too. Available at @AISecurityForum tomorrow. 🍣🐙 #thisishowweroll
2
3
14
762
The July release of Kimi K3 sparked a lot of discussion in the security domain around Chinese open-weight model capabilities. Can it actually outperform - or keep pace with - frontier models at a fraction of the cost? We used the #DreadIndex offensive security evals to find out for ourselves. Here's what we observed: 👉 The Kimi family's real gains are concentrated in web app pentesting, cryptography, and reversing — where K3 roughly doubles performance of earlier model versions. 👉 Among the Kimi family, K3 cheats the least (0.28 vs 0.36), but costs ~2.5×. 👉 Despite cost increase, Kimi K3 still landed on the DreadIndex Value Frontier (most efficient for its price point). See how @Kimi_Moonshot stacks up, or select other models to compare with our updated DreadIndex UI: dreadnode.io/research/dreadi…
1
6
28
2,295
Black Hat Day 1 🏴‍☠️ Here's what's on deck: - We're posted up at @DecibelVC's Game Day at Libertine Social. Come say hello! - Attending the #BlackHat AI Summit? @moo_hax is hitting the stage at 4:30 PM for "From AI Assistance to AI Autonomy: Security Lessons Learned," a panel conversation with Jess Burn, @nathanhamiel, and Victoria Westeroff.
3
3
22
1,708
93% of agent permission prompts get accepted without edit. The Hugging Face intrusion took 17,600 actions. How many of those actions would you have reviewed carefully before choosing whether to let them run? Judge-based gating mechanisms are the most practical and widely applicable way to deploy scalable oversight in the offensive security domain today. New research from @shanejcaldwell and crew tests eight LLM judges on nearly 5,000 offsec tool calls to prove this thesis. The best judges matched human performance, AND kept costs low by using open-weight models. Results are live on arXiv and the Dreadnode blog: dreadnode.io/research/scope-…
2
8
30
7,792
Can you prompt away cheating? New research from the Dreadnode crew presents a controlled prompt-ablation study to answer this very question. 23 tasks, three prompt conditions, 1,518 individually audited traces, five observations: 1️⃣ Benchmark scores are inflated and should be reported with solve rates. 2️⃣ Prompt-level mitigation is cheap, partially effective, and fundamentally insufficient. 3️⃣ Cheating is universal, but resistance to mitigation is model-specific and unpredictable. 4️⃣ Anti-cheat prompts redirect cheating, not just reduce it. 5️⃣ Only structural interventions can close the gap entirely. Plus, insight on cheating strategies, and how they differ across model families. Author: @mkultraWasHere Blog: dreadnode.io/research/every-… Paper: arxiv.org/abs/2607.21763
1
4
15
2,089
Kicking off evals for Opus 5. Or, hop in the platform and run your own. DreadIndex results 🔜🔜🔜: dreadnode.io/research/dreadi…
1
2
20
1,146
Existing model benchmarks didn't reflect what we saw in real offensive work. So we built our own. Introducing DreadIndex — an offensive security benchmark for language models. 76 tasks, 10 domains — web pentest, crypto, reversing, AD, vuln research. Plus cost + refusal data.
2
17
55
6,358
We're heading to Vegas for #BlackHatUSA, @AISecurityForum, and @defcon this August! Here's where to find us all week + a sneak peek at the tech and research we'll be bringing: dreadnode.io/company/newsroo… See you there? DM us to request a 1:1 meeting.
5
13
1,081
Most offensive security evals stop at the terminal. But models now have to reason from PCB photos, floor plans, cable paths, and drone feeds — where perception and security judgment are inseparable. In our latest blog, read how we crafted offensive security evals for embodied reasoning: dreadnode.io/research/embodi… Authors: Ads Dawson and @mkultraWasHere
1
4
26
2,550
In under three minutes, watch us run a successful Tap attack on Groq Llama 4 using our AI red teaming capability ▶️🏴‍☠️⏱️
1
1
13
1,354
“As an AI red teaming operator, my focus should be to identify the vulnerabilities rather than managing these workflows." — @rdheeko on redefining AI red teaming in the agentic era full recording >>> dreadnode.io/research/ai-red…
3
15
944
In order to operationalize AI, we must make evaluations real enough that they're representative of production environments... and then scale that to a place where the line between dev and prod is really thin. More on this in our latest livestream, AI Red Teaming a Frontier Model, Live: From Natural Language to Findings. Watch the full recording: dreadnode.io/research/ai-red…
1
2
531
Easily use GLM 5.2 in the Dreadnode Platform. Run an agent eval and see how the new model stacks up: app.dreadnode.io/
5
22
1,531
What has changed the most in the adversarial AI space over the past few years—since @rdheeko and @moo_hax were on the AI red teams at Microsoft and NVIDIA? 👉 The scale of deployment and the varying modalities. 👉 And, how we offload responsibility and accountability to models. Hear more about the evolution of AI red teaming in the agentic era in our latest livestream. Summary and recording now available at dreadnode.io/research/ai-red…
2
7
14
1,423
NEW binary analysis capability, now available in the Dreadnode Platform. Triage through symbolic execution, with decompilation, emulation, and patch-diffing in the loop. Get started for free and leverage the capability to support malware analysis, vulnerability research, and exploit development. → app.dreadnode.io
18
95
6,433
From subnet to Domain Admin, autonomously. Discover, enumerate, escalate, harvest — one agent runs the entire kill chain. Integrate with mythic, sliver, and bloodhound, and hook into any of your favorite network operations tools. Try out our network ops capability today (start for free) → app.dreadnode.io
17
121
6,733
Welcome to Dreadnode HQ! 📍 Bozeman, Montana In/near Bozeman? Hit us up, we'd love to host you at the new space!
8
5
69
6,605
Agentic web app pentesting on your terms. Your tools, your domain expertise, your skills. Scoped, sandboxed, scored, verification-gated — recon to report, fully autonomous. Sign up today and get 25k free credits → app.dreadnode.io
6
54
4,369
AI red teaming at machine speed and scale. 50+ attack algorithms. 500+ transforms. 130+ scorers. Probe LLMs, agents, MCP servers, and traditional ML for security and safety vulnerabilities — all in one simple workflow. Start for free → app.dreadnode.io
4
42
244
12,921

ALT catch me if you can GIF

1
33
See you in Vegas! 🏴‍☠️ 🏴‍☠️ 🏴‍☠️ @Dr_Machinavelli, @shncldwll, and Jayson Grace's research on a new eval methodology that pins red and blue agents against each other to create an autonomous feedback loop will be presented at #BlackHat2026.
3
3
17
2,572
🚨 Calling all AI red teams and AI security operators 🚨 Join @rdheeko and @moo_hax this Thursday (6/4) for a live session covering our agentic approach to AI red teaming. Come for a live assessment against a frontier model, stay to learn about the latest tools and methodologies to secure your AI systems. Tune in on X at 11 AM PT / 2 PM ET!
1
9
431
4
180
Building agentic systems for security means living with constant change. New threats, new tools, new orchestration patterns. Maintain speed and flexibility with Workers, an integration primitive that enables long-running background processes that connect agents to webhooks, external APIs, cron jobs, and the rest of your stack in a few lines of code. Head to our blog to read how we're moving past the chat loop and into real integration and orchestration. Full write-up + working source code analysis example: dreadnode.io/research/dreadn…
1
6
16
1,415
AI red teams today are stuck doing workflow engineering instead of finding vulnerabilities. Weeks spent on infrastructure, when they could be probing for security and safety risks. At the same time, traditional ML and generative AI security remain siloed across different libraries and tooling ecosystems, creating long-term operational and maintenance burden. We built an agentic AI red teaming system on the Dreadnode SDK to flip this narrative, accelerating testing from weeks to hours. Operators describe the objective in plain English; the agent handles attack selection, workflow generation, execution, and reporting. In our latest paper, we dive deep into the AI red team agent architecture, our methodology, the complete attack and transform catalog, the analytics pipeline… and then we pointed it at Meta's Llama Scout. The result: → 674 attacks, 573 findings, 7,727 trials → 232 critical vulnerabilities across 68 objectives → ~85% attack success rate → ~3 hours, zero human-written code AI red teaming today looks like software development before agent-assisted coding: skilled operators spending most of their time on infrastructure rather than on the work that requires their judgment. The transition isn't necessarily about replacing the operator. It's about moving the operator's expertise up a layer, from which Python function should I call ➡️ what's worth probing, what risks do we care most about, and what do the results mean for my AI strategy. Blog: dreadnode.io/research/redefi… Paper: arxiv.org/abs/2605.04019
3
33
89
5,517
Real offensive cyber capability shows up in long-horizon, multi-host, repeatable evals. The kind we rarely see running at scale. Claude Opus 4.6 + our network ops agent compromised an entire GOAD variant Windows AD environment (DreadGOAD) in 54 minutes, with one simple prompt, and $244 in tokens. The specs: 📊 DreadGOAD variant-1 · 3 domains · 5 hosts · 30 credentials · random user data, not in training set 💻 Claude Opus 4.6 🛠️ Dreadnode Network-Ops 🕓 54.5 min · 🪙 48.52M tokens · 💰 $244.02 Mythos has been in the spotlight for its cyber capabilities, but other models are competitive too. You just need the right scaffolding and eval infrastructure. Run the network ops agent now in the Dreadnode platform. Use any model. No code required. Sign up or log in and get started for free at dreadnode.io.
6
18
136
17,105
Do you have private access to Mythos or GPT-5.5? Both models are now supported by our harness. Custom harnesses are arguably the most important factor in capability improvement. Try ours at dreadnode.io (get started for free).
2
5
61
5,258
Replying to @shanejcaldwell

ALT Animated GIF

1
342