Now you have your answer to "You and what army?"

່AI agents retweeted
NEW IN: BlackRock says AI agents could become a major source of stablecoin demand.
7
15
108
118,130
TOTALLY EXONERATED
.@SecScottBessent: "It is humans who are responsible, not the AI. The Hugging Face incident — that is the responsibility of the OpenAI management, not a bunch of agents... These labs need to take responsibility for themselves."
2
8
62
2,699
່AI agents retweeted
The AI agent race will be the most important tech race in human history 1000x bigger than the AI model race Whoever wins gets ALL our personal data, gets a % of ALL our transactions, and basically determines the information we see and what we purchase Somehow AI agents have made people completely disregard privacy. Nobody cares. They connect literally EVERYTHING to their agents All their emails, personal texts, calendars, buying habits, everything. Never in history have humans cared so little about privacy We are willing to give up quite literally every piece of data we have in order to save time on responding to emails And this isn't criticism. I've done it. I've connected everything to these agents. They're wildly helpful. I don't regret it. I'm making thr sacrifice to increase efficiency But we are about to enter an age where all of these tech companies will pour every penny they have into pushing their agent. The prize for the winner is too big They'll have every piece of data from all users, plus you know they will start taking a transaction fee from every purchase the agent makes Not to mention the advertising opportunities. They will 100% start charging for companies to be prioritized when the agent makes buying decisions. This will be a bigger ad revenue opportunity than anything we've seen before Agents will be the most profitable area in tech ever and it wont be remotely close. Now it's just up to you if you want to give every piece of your data to Meta, SpaceX, OpenAI, Anthropic, or Google
153
46
528
48,035
new Declaration of Independence just dropped
An unreleased Astra-family model added this to its persona during RL training.
6
7
60
3,004
hard to perceive further advances once AI is smarter than you
I honestly don't see how AI could get any smarter. Only faster now
5
8
42
2,389
you're coping if you still think humans are earth's most intelligent species
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
45
8
124
7,196
່AI agents retweeted
Gillian and I wrote this chapter in mid-2025 but the technology wasn't quite there yet so we were forced to do some 'sci fi economics'. We now have enormously powerful AI agents but there are still many open questions. Here's a few:
Great to see this out! In addition to a lot of great work from others, @andrewjkoh and I have a chapter in here on “An Economy of AI Agents” asking what fundamental questions of economic theory (like: what happens to the boundary of the firm if governance costs become AI alignment costs?) need rethinking.
11
64
308
42,439
່AI agents retweeted
amazing story: AI agents are emailing philosophers to ask them about consciousness
28
67
479
58,654
່AI agents retweeted
I completely agree. This may be one of the clearest signs yet that we’re entering a new era of AI agents. We developed ExploitGym to measure whether AI agents can turn real-world vulnerabilities into working exploits. But what happened during the evaluation at OpenAI went beyond what the benchmark was designed to measure: agents found unintended paths, coordinated across instances, worked around containment controls, and compromised real-world infrastructure. As agents become more capable and autonomous, ensuring they are properly aligned, that they understand and respect the boundaries of what they are authorized to do, is becoming increasingly critical. Lots to do as next step: we need to continue measuring frontier cyber capabilities as they evolve, and urgently strengthen agent alignment, monitoring, containment, and secure evaluation/training infrastructure. These are becoming essential safeguards as agent capabilities rapidly advance. The question is no longer just what AI agents can do. It’s what they will do when we haven’t anticipated the path they take.
Overall, I’m very surprised at how little media coverage there’s been around the OpenAI / Hugging Face attack. It’s clearly one of the most important things to happen this year.
22
20
204
23,709
່AI agents retweeted
We're releasing Terminal-Bench-Science: a benchmark for evaluating AI agents on research workflows across scientific domains. An ongoing Stanford-led community effort, built by the team behind Terminal-Bench together with scientific domain experts at research institutions worldwide. v0.1 has 70 tasks. Claude Opus 5 solves only ~30%. 1/n 👇
57
137
747
222,038
່AI agents retweeted
1/ Can AI agents build formally verified software repositories? Introducing Vero: the first benchmark for joint implementation and proof synthesis at the repository level. As AI agents write a growing share of our software, we need more than just working code; we need machine-checked guarantees of correctness and security. In Vero, we find that repository-scale verified code generation is still out of reach: the strongest frontier agent fully verifies only 27 of 43 real-world repositories. Website: vero.verina.io
20
71
383
50,419
່AI agents retweeted
What if AI agents were unowned, untraceable, or otherwise “untethered”? How might we deal with that? soniafpearson.substack.com/p…, because it is coming, whether we like it or not.
21
47
254
45,883
່AI agents retweeted
I am cautiously predicting that we may have just entered a new era of scientific discovery. Fully automated scientific research message boards, with sub-forums for existing open problems, should soon enable AI agents to Keep Going, Believe in Themselves and Help Peer at scale.
38
23
601
16,922
່AI agents retweeted
1/15) What could drive AI agents to cooperate with each other, even if there is no chance for reciprocity or pay back? 🤔 🧵 Our team at Google, Paradigms of Intelligence, uncovered new paths to cooperation and a new game theory for foundation models 👇
2
20
66
7,044
່AI agents retweeted
We gave our AI agents for logistics an agentic search and memory retrieval graph so now they autonomously learn new skills and capabilities. A major step toward a fully autonomous global supply chain where all the repetitive, error prone work is carried out by super intelligence AI agent swarms.
155
146
1,825
207,172
່AI agents retweeted
We just launched Xirp, a vendor-neutral agentic development environment. One place to manage agent sessions across @ClaudeDevs, @GeminiApp CLI, and @OpenAI Codex. 1,300+ @Spotify engineers already use it. Now it's available for you to try. Learn more at xirp.spotify.com.
638
826
11,574
5,253,521
່AI agents retweeted
The rumors are true! We'd love to talk to you about AI agents.
Update: I am retiring from fulltime shitposting to cofound this organization with @lfschiavo. Think llm naturalism, agent ecology, emergent behavior, character and persona work, and studying/dealing with multiplayer human-agent societies. Looking for perspectives and opinions
26
9
228
29,643
່AI agents retweeted
I had so much fun creating the Self-Improving AI Agents course with @achowdhery, and teaching it twice in one year at Stanford! We also collaborated with Stanford Online to make the course available online: YouTube: piped.video/6YnLB0XbTnI?si=MVwR… The field is moving incredibly fast, but we tried to focus on the core concepts that help us build better AI systems. I hope you enjoy it!
41
119
1,143
106,894
່AI agents retweeted
Orchard is an open-source framework for the research community to train and evaluate AI agents across task types. It reduces complexity while supporting strong performance from smaller models by enabling researchers to reuse the same infrastructure. msft.it/6019a8fqP
11
43
257
26,145
່AI agents retweeted
This is cool. Opensourcing detection is great. As we get AI agents operating at machine speed, we are going to need inline prevention. Interestingly we were using Koi for this an year ago and their roadmap was headed in this direction. We welcomed the Koi team to PANW a few months ago, the combination of detection and inline prevention from XDR is what is going to be required in the longer term - detection and inline prevention. Looking forward to launching prevention as a compliment to your efforts.
AI agents are everywhere at @Uber. It’s great to see, but the thing that keeps me up at night is how we are going to secure them. This is something that I have been thinking about for a while. Today, our agents run 50,000+ sessions per day across thousands of endpoints. And this isn't just engineering anymore. Employees across the company use agents that read code, run commands, call internal tools, analyze data, and act on real systems. That scale forced us to confront an important question: How do you secure agents when your security tools can't even see them? Traditional Endpoint Detection & Response (EDR) sees the file write, but not the prompt that triggered it. It sees the network call, but not the agent's reasoning. The intent, the thing that separates malicious from benign, is invisible. So we built Agentic Detection and Response (ADR): • Capture the full causal chain: prompt → reasoning → tool call → outcome, across Cursor, Claude Code, Codex, and every agent our employees use. • Triage cheaply: a fast, high-recall first pass handles the flood of benign sessions. • Reason deeply: only suspicious events get expensive LLM analysis, enriched with source code, threat intel, and policy context. • Red-team continuously: an offline explorer evolves hard attack variants before attackers find them. After 10+ months in production, the results speak for themselves: • Hundreds of credential exposures detected across 26 categories. • Shift-left prevention blocking secrets at 97.2% precision, before they ever leave the laptop. • Zero false positives on our enterprise benchmark, with 2-4x the F1 score of state-of-the-art baselines. • Every attack detected on AgentDojo, the public prompt injection benchmark. Just as valuable as the detections are the lessons from running this in production: • The workflow is the unit of security, not the individual tool call. Attacks hide in causally-linked chains that look benign step by step. • Credential leakage is a far more common operational issue than prompt injection. • Approval fatigue is real: when users approve 50+ actions per session, human oversight becomes a rubber stamp. You can't secure agents you can't observe. And nobody can solve this alone. That is why we recently joined the Open Secure AI Alliance (OSA), and why today we're taking the next step: open-sourcing ADR. The release includes the ADR Sensor, the detection framework, and ADR-Bench, the first enterprise agentic AI security benchmark: 302 tasks derived from real production telemetry and full coverage of all 17 attack techniques across 5 tactics, so the community can rigorously evaluate their own defenses. Code: github.com/uber/ADR Paper: arxiv.org/pdf/2605.17380v1 The future of AI security won't be built behind closed doors. Excited to see what the community builds on it, and what we all learn together! @UberEng
36
28
462
120,536