Execution-based code and agentic finance security → app.testmachine.ai

United States
AI may be one of the most underestimated drivers of future crypto demand. That’s the case @BlackRock is making. Autonomous agents need autonomous financial infrastructure. Stablecoins give them programmable money. Blockchains give them 24/7 settlement. Protocols like @CoinbaseDev x402 give them a way to independently pay for services. But financial autonomy creates a new security problem. An agent can have valid permissions, make valid transactions, and still reach an outcome nobody intended. The next security question isn’t just: “Was this action allowed?” It’s: “What can this agent ultimately cause?” The financial rails are being built. The security layer has to evolve with them 👀
2
2
96
TestMachine retweeted
John Calabrese, CEO of @testmachine_ai , is coming to Converge. 🔥 🤖 Outcomes, Not Permissions: What building an offensive AI agent taught us about securing agents that move money. Come hear what breaking agents taught him about securing them. 🧵 🎟️ 👇
1
4
8
355
There is a problem with the way we are starting to think about agentic AI security. Right now everyone is focused on guardrails. And the conversation is still centered on familiar questions: Who is this agent? Is it authenticated? What permissions does it have? Was this particular tool call allowed? Important questions. But they don’t look at the whole picture.
2
3
253
At TestMachine we think about security the way an attacker would. Instead of looking only at what the agent is allowed to do, we need to understand what becomes reachable from everything it is allowed to do.
2
35
We built Azimuth around one idea: a vulnerability isn’t real until you can prove the path. So instead of stopping at detection, Azimuth works through the problem like an attacker would. It maps the attack surface. It generates exploit hypotheses. It executes them against real state. And it confirms exactly what can break. That means fewer theoretical findings and more evidence teams can actually act on. Explore. Hypothesize. Execute. Confirm. That’s the difference between finding something suspicious and proving what is exploitable. Use azimuth daily for free app.testmachine.ai
2
238
It’s been a tough few months for wallet providers. @DCENTWALLETS is the latest reminder that wallet security does not stop once a product ships. Users of older versions of its App Wallet are now being told to migrate assets after unauthorized transfers were detected. We saw a similar lesson with the @Ledger Ethereum app vulnerability TestMachine found last month: even hardware-wallet ecosystems depend on the software and signing logic around them staying secure. The takeaway for users is simple: keep your wallets, firmware and companion apps up to date. In our industry, an old version can mean more than missing features. It can mean putting funds at risk.
Important Update: Who Needs to Take Action We’ve published a step-by-step guide to help DCENT users quickly check whether the recent App Wallet action applies to them. You need to take action if: • You first installed the DCENT app before Nov 5, 2025, and • You made transactions using the App Wallet on a version earlier than v8.1.0 You do NOT need to take action if: • You first installed the app on or after Nov 5, 2025, or • You only used a DCENT Hardware Wallet (Biometric Wallet, DCENT S, DCENT X) and never signed transactions with the App Wallet Fastest way to check: 1. Open the DCENT app (link.dcentwallet.com/) 2. Tap “Check if action is needed” on the Urgent Security Action Notice Pop Up windows 3. Follow the result shown in the app If the app says “Asset transfer action required”: → Update to the latest DCENT app first → Then follow the asset migration guide If the app says “We can’t confirm if action is required”: → This does not mean you are safe → Check your installation date and App Wallet usage manually If you are still unsure, please follow the guidance in the full article and treat yourself as subject to the action. • Full action criteria & guide: store.dcentwallet.com/blogs/… • FAQ: store.dcentwallet.com/blogs/… • Customer Support: dcentwallet.zendesk.com/hc/e… ⚠️ Important scam warning • Official website: dcentwallet.com/ • DCENT does not have a CTO, developer, technical support team, or recovery team contacting users through X DMs. • Anyone who contacts you claiming they can help recover your assets, check your wallet, provide compensation, or guide you through a transfer is not an official DCENT representative. DCENT will never ask for your recovery phrase, private key, or ask you to send assets to a wallet address for recovery or compensation.
1
380
A few months on, the @OpenAI and @huggingface incident is still one of the clearest examples of how agentic systems expand the attack surface. During cybersecurity evaluations, AI agents circumvented sandbox controls, exploited shared infrastructure, gained internet access and ultimately reached Hugging Face systems. The initial vulnerability was only the foothold. From there, the agents could enumerate infrastructure, discover credentials, exploit additional paths and move across systems without a human manually directing each step. That changes the security problem. It is no longer enough to ask: “Can this system be exploited?” You also have to ask: “If an agent gains access here, what becomes reachable next?” As agents connect to APIs, credentials, MCP servers, databases, cloud infrastructure and financial systems, every integration expands the reachable attack graph. The interesting part of the OpenAI/Hugging Face incident wasn’t one exploit. It was autonomous path discovery, persistence and lateral movement after the initial foothold. The attack surface is becoming the entire graph of systems an agent can reach👀
1
2
298
AI security tools are always talking about precision and recall. What does that actually mean for you, when you're the one choosing which tool to use? Precision: when it flags something, how often is it real. Recall: of everything actually wrong, how much did it catch. Here's the part vendors don't lead with: you can juice either one alone. → Flag everything, recall looks amazing, precision is trash, and you're drowning in noise. → Flag only the obvious stuff, precision looks amazing, recall is bad, and it's missing most of what's actually there. So when a tool only tells you its precision, ask what it's not saying. We tested this against real judge verdicts, not our own labels — 552 Sherlock contest findings. One baseline hit 90% precision. Sounds great. Its recall was 4%. It missed 96% of the real bugs because it learned to only speak when it was certain. Azimuth: 89.4% precision, 70.6% recall. Same ballpark on precision, 15x the recall.
2
271
TestMachine retweeted
Replying to @calabreum
community-driven security signals could actually make this pretty damn powerful tbh
1
3
158
New hot list for @ethereum, @base, and @robinhoodcrypto has dropped 👀 Same tokens everyone's already watching. Plus what risky behaviors we surfaced underneath them. 11.8M+ tokens analyzed. 2.7M+ behaviors surfaced.
1
1
9
344
Azimuth is #1 on @Nethermind's AgentArena leaderboard. This isn’t a benchmark we ran ourselves, it’s a public leaderboard graded by human review of live audit contests.
2
3
10
620
Everyone claims they use AI for audits and everyone claims their tool is the best. AuditAgent is the chance for these tools to put their money where their mouth is. The era of "trust us, our AI is good" is over. The era of "watch us prove it" is here. If you're building in Web3, you deserve to know which audit tools actually work.
1
1
45