Agentic security to outpace AI attacks. AI agents that continuously finds, exploits and proves real vulnerabilities in your systems. #1 on CVE bench.

Secure your system →
Announcing FailSafe's listing as a solution provider on the @MAS_sg Pathfinder Programme (PathFin.ai). PathFin.ai is a collaborative initiative between the Monetary Authority of Singapore and the financial industry that fosters knowledge exchange in Artificial Intelligence (AI) implementations. For financial institutions in Singapore, this means FailSafe’s agentic offensive security platform is now among the solutions available for consideration through the programme. FailSafe SWARM continuously pentests applications, APIs, cloud and AI systems the way a real attacker would. Confirmed findings include proof-of-concept evidence, remediation guidance and audit-ready reporting, supported by traces showing how vulnerabilities were identified and validated. The practical difference is cadence. Offensive testing that keeps pace with release cycles rather than arriving once a year. SWARM achieved the highest publicly reported score on CVE-Bench v2.1.0, an autonomous black-box exploitation benchmark, scoring 62.5% in the zero-day setting on a single run per target. Let's chat about a new security model for the age of AI.
1
2
2
380
this is the message
In 100 days, your board will expect a response to agentic threats. In 100 days, you're gonna be facing roughly 10x the number of agents probing your attack surface. Are you ready?
2
397
FailSafe retweeted
Announcing FailSafe's listing as a solution provider on the @MAS_sg Pathfinder Programme (PathFin.ai). PathFin.ai is a collaborative initiative between the Monetary Authority of Singapore and the financial industry that fosters knowledge exchange in Artificial Intelligence (AI) implementations. For financial institutions in Singapore, this means FailSafe’s agentic offensive security platform is now among the solutions available for consideration through the programme. FailSafe SWARM continuously pentests applications, APIs, cloud and AI systems the way a real attacker would. Confirmed findings include proof-of-concept evidence, remediation guidance and audit-ready reporting, supported by traces showing how vulnerabilities were identified and validated. The practical difference is cadence. Offensive testing that keeps pace with release cycles rather than arriving once a year. SWARM achieved the highest publicly reported score on CVE-Bench v2.1.0, an autonomous black-box exploitation benchmark, scoring 62.5% in the zero-day setting on a single run per target. Let's chat about a new security model for the age of AI.
1
2
2
380
FailSafe retweeted
Shipping on BNB Chain? Security support just got easier to find. The AvengerDAO marketplace brings together 11 security teams covering audits, threat monitoring, risk scanning and incident response: @HashDit @CertiK @zokyo_io @salus_sec @Beosin_com @GoPlusSecurity @BlockSecTeam @pessimistic_io @sherlockdefi @getfailsafe @FirepanHQ Explore the marketplace: avengerdao.org/marketplace Learn how the AvengerDAO marketplace works in our blog 👇 bnbchain.org/en/blog/avenger…
34
24
152
52,316
FailSafe retweeted
GLM-5.3 API is now live. - Built for coding, defensive cybersecurity, and long-horizon agentic tasks - Priced the same as GLM-5.2 - Available via the official API and partner model gateways Get started: docs.z.ai/guides/llm/glm-5.3
189
327
3,783
528,216
FailSafe retweeted
FailSafe SWARM is #1 on CVE-Bench. 62.5%. Nearly double the next best reported result.
2
13
223
8,842
FailSafe SWARM is #1 on CVE-Bench. 62.5%. Nearly double the next best reported result.
2
13
223
8,842
If you're building in this space: post your CVE-Bench score. The benchmark is free, open, and the bar is public now.
58
+1. And if you’re a company who cares about security but don’t know where to start, dm us. We believe that the only way to defend against machine level intelligence is to arm the good guys with the same level of capabilities.
People are not appreciating one of the biggest takeaways from the COLDCARD hack. Cybersecurity is now all about spend. Once AIs are doing all of the attacking, the simple question is how much money are you spending with frontier AIs scanning for vulnerabilities, compared to what attackers are spending? The fact that Claude was able to find this in 8 minutes (and GLM in 20 minutes) tells you that these guys were doing basically nothing. Absolutely 0 AI hardening happening on releases. COLDCARD is a minority vendor. They have probably ~2% market share within Bitcoin hardware wallets. The lesson: security-sensitive products (crypto wallets especially) are going to consolidate toward larger, better-funded players who can afford to do the security hardening, and small companies without meaningful funding are simply NGMI. Second: we can now start standardizing reporting on the cost of discovering an attack, in order to understand how bad it was. Call it Cost of Discovery (CoD): how much it would cost a frontier model to discover the same vulnerability. The linked Claude Code claim is a little suspect (it may have been contaminated by web search), so another poster turned off web access to GLM 5.2 and was able to rediscover the attack in 20 minutes. Taking the @ZhipuAI API numbers ($1.40/M in, $4.40/M out, 35tps), given the kind of workload here (read-heavy agentic workload), Opus estimates the total Cost of Discovery on this bug was around $2. $2 of AI hardening would've caught this bug. There is no excuse for this. In the future we should start reporting Cost of Discovery on these when vulnerabilities are independently replicated. (Companies should not necessarily publish the amount of money they are spending on AI-hardening, as that would imply to an attacker: if you spend more than X, you may find something.) If you are a startup and building anything, you should be running AI-hardening on EVERY release. You should be spending AT LEAST in the thousands of dollars using a frontier-level model searching for crits, especially when the endpoint (recovering the key, or draining money out of a smart contract) is easily verifiable. My heart goes out to anyone who lost money from this. AI is genuinely changing the game and many have not been able to adapt in time. But I do believe in the long run, AI is going to make us all better off security-wise than before. We have no choice but to adapt.
1
6
839
FailSafe retweeted
FailSafe Swarm has been cooking with @nvidia's Nemoclaw, @NEARProtocol 's Openclaw, @NousResearch's Hermes and more. Keeping the world safer, one PR at a time. receipts below 👇🧵
2
5
13
3,379
Attending @superai_conf ? Come for the closing event of the week, where our cofounder @0xWui will be chatting about agents, enterprise deployments and more! luma.com/mqxgtrj5
1
1
10
484
FailSafe retweeted
@getfailsafe is hosting an enterprise agent mixer with @awscloud , @coinbase , @NEARProtocol and friends. Agents are moving from demos to production. That means new questions around evals, security, runtime controls, and trust. We’ve been spending a lot more time in the Bay lately as we deepen our R&D presence here. Come join us!
2
2
8
542
If you’re on GitHub - stay safe! Enforce MFA if you have not!
We are investigating unauthorized access to GitHub’s internal repositories. While we currently have no evidence of impact to customer information stored outside of GitHub’s internal repositories (such as our customers’ enterprises, organizations, and repositories), we are closely monitoring our infrastructure for follow-on activity.
1
3
589
Had a great time hosting ai-pilled builders and enterprise leaders with @okx and @awscloud last week. AI agents are transacting onchain and @getfailsafe is securing them across every layer of the stack.
4
2
10
668