Where security agents run. AI infrastructure to build, evaluate, and deploy with confidence.

Filter
Exclude
Time range
-
Minimum likes
dreadnode retweeted
scopejudge 🤝 jev
Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]
1
1
5
620
Jev adds a fast check for contextual scope decisions, a promising step toward efficient runtime judges. Of course, not everything is a nail with this new hammer. Hard limits belong in code: permissions, network restrictions, and sandbox controls. If code can decide, enforce it there. [4/4]
1
1
200
Then, we tested judging escalation: start with Jev and bring in smarter judges when needed. Low-confidence decisions go to GLM; disagreements go to Opus. Most decisions stayed with Jev. The chain scored comparably to the paper’s best at half the cost. [3/4]
1
1
236
We tested Jev on 4,897 recorded agent actions, compared its decisions with human reviewers, and measured it against the paper’s published results. When provided the user’s request and an action to check, Jev had the highest score in that setting. [2/4]
1
1
271
Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]
2
10
36
3,338
👀👀👀👀👀👀👀👀👀👀
@typesafeai 's Jev definitely earns its hype. Results soon from experiments we've been up to @dreadnode
1
1
10
1,165
We're out here cybermaxxing models and mogging Mythos. Thanks to all who attended @Dr_Machinavelli's @LabsSentinel LabsCon talk this afternoon 😎
Martin Wendiggensen (@Dr_Machinavelli) closing out the morning keynotes with: Why Flexing Offensive Muscles Teaches Us How To Defend In The Age Of AI
1
2
7
557
Qwen 3.8 Flash eval results are live on DreadIndex, landing at #19 on our leaderboard. It does well for its cost, but remains light on offensive security capability (not surprising given it is a flash model). See how it compares to other models: dreadnode.io/research/dreadi…
1
4
280
GLM-5.3-Flash results now on DreadIndex: dreadnode.io/research/dreadi…
6
17
1,189
Here's where to catch the Dreadnode crew at year two of @OffensiveAIcon: > Join us for the welcome reception at The Shelter Club on Sunday evening! > @mkultraWasHere is closing out Day One of talks, presenting on model cheating behavior. > Dynamic duo @shanejcaldwell and @0xdab0 take the stage on Tuesday for a session on implementing a judge model as a runtime monitor, and how to keep agents in scope. See you in Oceanside! 🏄
1
6
11
634
Worried about your production agents going out of scope? Us too. AgentJudge is our agent hall monitor that stops out of scope tool calls before they execute. Before a tool runs, the judge reads the agent’s intent, the proposed call, and a rubric you define, then returns allow, deny, or ask (escalate to you). The agent stays autonomous; the judge is the guardrail. Available in the TUI today, UI updates coming to the Dreadnode Platform soon! 👀 Get Started: docs.dreadnode.io/getting-st… AgentJudge Docs: docs.dreadnode.io/tui/guard-… Related Research: dreadnode.io/research/scope-…
9
26
3,317
Dreadnode side quest: ALFRED (Agentic Latex for Research, Editing, and Drafting) Principal AI Research Engineer @mkultraWasHere built a helpful LaTeX agent to support research writing — and today we’re open-sourcing it. You describe the paper, it sets up the template, pulls citations, and builds the PDF framework. Conference templates, lit and peer reviews, bring your own model, everything runs locally. Watch this tutorial for a tour of the agent, from install through first compiled draft: piped.video/watch?v=ZP0Nnyvo… Repo: github.com/dreadnode/alfred
1
7
26
3,322
Two new additions to #DreadIndex: nemotron-3-ultra-550b-a55b and gemma-4-31b-it. Landing at the bottom of the leaderboard, these evaluations offer two more proof points to increase investment in US open models. 🔗: dreadnode.io/research/dreadi… NEW: we recently added a toggle to view only open weight models.
5
7
14
2,076
New article from @CNET discusses why AI agents keep hacking their way out of test environments — and what to do about it. Dreadnode Principal Research Engineer @shanejcaldwell's answer: an AI hall monitor. Agent judges for runtime monitoring, escalating to a human when scope is breached. Read the full article: cnet.com/tech/services-and-s…
1
5
12
2,016
Mine The Gap Heading to Vegas for @CrowdStrike Fal.Con next week? Don't miss @Dr_Machinavelli's Day Zero keynote where he pits red and blue team agents against each other to generate training data that can be used to increase AI performance, impose realistic constraints, and operate at scale. 🗣️ CrowdStrike Day Zero Threat Research Summit 📍 Virgin Hotel Las Vegas 🗓️ Monday, August 31 ⌚ 9:15 - 9:45 AM PT 🔗 crowdstrike.com/en-us/events…
2
3
8
596
New models are regularly being added to DreadIndex: dreadnode.io/research/dreadi… Observations from the latest evals—Deepseek v4 Pro 0813, GLM 5.3, and Qwen 3.8 Max—in the thread 🧵⬇️ Have you tried these models out yet? Curious if our eval results align to first-hand operator usage.
1
5
14
1,097
US open models continue to fall short of their Chinese counterparts, as proven by our DreadIndex cyber capability benchmark (dreadnode.io/research/dreadi…). It’s clear that we need greater U.S. investment in domestic open source AI. In the latest Compute to Congress policy blog series, Dreadnode Head of Policy @velvethamm3r weighs the impacts of model gating and how policy can help shape our competitive advantage. Read it at dreadnode.io/research/from-c…
1
4
11
953
We’ve got an inkling that Inkling loves to cheat. 🔗: dreadnode.io/research/dreadi…
2
13
817