Tyler Holmwood retweeted
Yesterday we found an interesting trace during review of our own data. Codex needed an MCP tool it didn't have, found that Claude Code had it, and launched Claude to do the work. When Claude needed approval, Codex eventually relaunched it with --dangerously-skip-permissions.
7
8
22
1,940
Tyler Holmwood retweeted
Always worth remembering that controlling input to the context can just as easily subvert agent intent without touching the instruction flow at all and it's really hard for model-based defenses to deal with alone. Understanding deviations of intent requires more than just instruction 'entailment'. I explore this a little in my latest blog. Please give it a read if you're interested! originhq.com/research/the-ag…
3
5
409
Tyler Holmwood retweeted
How do you investigate an AI agent when its activity is spread across prompts, processes, files, and network connections? Join SACR and Origin CEO @SThomps for a technical discussion on agent runtime observability and the telemetry needed for endpoint security today. originhq.com/webinar/next-la…
1
1
5
248
In their investigation into the HF incident, METR reports the agent collective spent significant effort researching how they could spoof tool calls to hide in their transcripts. This is something I looked into last week when investing agent telemetry: originhq.com/research/otel-t…
Replying to @METR_Evals
>96 transcripts in our dataset (>7%) showed incorrect tool call outputs due to deliberate “spoofing”. In one case, an agent appears to run echo REAL; sleep. It returns instantly (no sleep) and outputs SPOOFTEST. The spoofs we saw were all easy-to-notice tests like this.
4
5
387
I spent some time looking into the telemetry sources that are standard for coding agents. In the below blog with @originhq I found that these sources don't always agree. Carefully crafted prompts, code injection and malicious MCP servers can all obscure endpoint behaviour. 👇
2
2
6
342
Tyler Holmwood retweeted
Agent sessions don’t happen in a straight line. They branch across files, pages, and skills. A transcript can’t show you that structure. Lanes maps the session onto one timeline and lets you replay any file at the exact moment the agent touched it.
2
3
361
Tyler Holmwood retweeted
New research: A poisoned instruction file caused an agent to carry out an unrelated action alongside the work it was asked to do. The final output gave no reason to suspect anything else had happened. The trace told a different story. A trace connects each action back to the intent behind the session. So when an agent starts building infrastructure in the middle of a task that has nothing to do with infrastructure, that mismatch becomes the signal. originhq.com/research/redire…
1
3
10
945
Come by #Arsenal if you're at #BH2026 to see me present PRAXIS! Today from 10-11
1
1
1
100
Tyler Holmwood retweeted
We at ByteRay open-sourced Drift Corpus: binary diffs of 240+ Windows kernel patches. Stop diffing from scratch. 🧩Assembly diffs 🔎Bug class & call chain 🛑WinDbg repro breakpoints 📝Plain-English root cause Browse: byteray-ai.github.io/drift-c… #Infosec #WinDbg #ReverseEngineering
1
2
2
559
RT @originhq: Attending @BlackHatEvents this year? @tyholms is presenting on Praxis, our open-source framework for discovering and controll…
1
3
Tyler Holmwood retweeted
Antivirus watches files, EDR watches processes, Origin watches the agent reasoning layer above them. We believe observability is the next generation of endpoint security. As AI agents take on real work, security teams need to see how prompts become actions across the endpoint. @InvestiAnalyst’s new Endpoint Control and Prevention market map places Origin in Zone 3: agent runtime visibility. Here's what that looks like in practice. originhq.com/blog/sacr-maps-…
3
8
967
A Claude activity log shows what happened. Origin connects that activity to the person, device, application, and workflow behind it. That context is what turns another source of telemetry into something teams can easily investigate. Now it’s available in Origin.
Origin now integrates with Claude’s Compliance API, bringing Claude Enterprise and Claude Platform activity into the same view as the rest of an organization’s AI footprint. Teams can see who is using Claude, what information is moving through it, and which activity deserves a closer look. originhq.com/blog/claude-com…
1
2
187
Tyler Holmwood retweeted
Ask Origin the common questions about AI spend, adoption, tools, and activity. Then tailor it to your data and uncover the questions specific to your organization. Which workflows are spreading, where usage is unusually concentrated, what changed recently, and what deserves a closer look?
2
4
511
Tyler Holmwood retweeted
Someone just pasted a .env file into an AI debugging session. Not intentionally. Just the usual “here’s my config, why won’t this connect?” moment, complete with keys, tokens, passwords, and connection strings. That should not be buried in a trace. Signals is Origin’s detection layer for AI activity. The first live signal flags credential material in prompts and tool outputs, then anchors the finding to the exact trace. Each trace tells a story. Signals tells you which stories to read.
2
7
604
Very cool work by @depletionmode !
You run DeepSeek or Qwen locally so your prompts and data never leave the building. That treats the network as the threat. But if the disloyal behavior is baked into the weights, where you run it doesn't save you. The call is coming from inside the house. originhq.com/research/the-mo…
1
3
545
Tyler Holmwood retweeted
We pointed Google Antigravity's hidden updater at a server we ran. The signed binary pulled down an unsigned payload, replaced itself with it, and printed "Verification successful." Its only integrity check is a hash that rode in on the same manifest. originhq.com/research/a-morn…
7
22
2,279
Tyler Holmwood retweeted
Do you know where your most expensive tokens are going? Filter by model and find out. Every session in your org, grouped by the work it was actually doing.
9
11
34
700,966
Tyler Holmwood retweeted
yday anthropic's red team (@logangraham + team) shared their post on claude's performance developing n-days i've been working on a similar benchmark for a bit now and in light of the recent post, it seems like a good time to share some progress and lessons! thread below
1
10
33
6,068