Harry Wetherald retweeted
Replying to @dash0hq
@dash0hq, @proximafusion, @WeAreLegora, @bfl_ai, @WithCoral_com, @Maze_Security, @EntireHQ, @ZaroAI_, Manifold, Forgent, Integral + some of our stealth 🍒's. Every one of them is f***ing crushing it. Amazing to see European founders setting the pace in AI. 🇪🇺🔥🍒
3
29
1,608
Very similar to how we think at Maze. It sounds cool to train your own model but it’s not what matters when you’re building multi-agent systems handling specialised tasks. Slightly tuning the weights of a model matters much less than training and orchestrating the entire system.
There is no best model. There's a lot of noise about models right now. Who is training them, who owns them, where legal intelligence should live. One question actually matters: what produces the best outcome for the legal task in front of you? That's how we decide things at @WeareLegora. We optimize for the end-to-end outcome on a legal task. The model is one layer of that system, not the system. Models are uneven and the frontier changes almost weekly. One model plans a long job well, another runs deep analysis across thousands of documents. Some have to be told exactly what to do, and some are fine with a vague brief. They all break in different ways. So our lawyers write evals and we test them with the Legora BAR, our benchmark for agentic reasoning. Every model takes every test, and the model that wins gets the work. We post-train when we know it buys our customers better performance on a specialized task. Training is a tool we reach for when it helps, nothing more than that. The intelligence that compounds sits in the orchestration layer. Precedents, review standards, client requirements. That knowledge has to stay editable, auditable and portable. In our system, a changed review standard is an edit that takes effect the same day, with no new model training required. No lawyer should have to worry about which model did the work, any more than they think about which chip is in their laptop. They should only care about the quality of the work. That's what we are focused on. If you want the engineering version of this argument rather than the CEO version, our CPO, Bryan Tsao, and CTO, @jacsebl, take it apart in the video below.
1
3
83
Tech leaders in the 2010: "Don't worry about social media! We're the good guys! We just believe in connecting humanity!" Tech leaders today: "Honestly, this thing we're building is probably going to kill everyone."
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
105
Harry Wetherald retweeted
The counterintuitive truth that most companies want to avoid: Right now the quality of your engineering team has never been more important
9
7
116
9,669
Anthropic: "Hypothetically, our models could maybe commit major cyber attacks." OpenAI this week:
1
57
Harry Wetherald retweeted
One team's AI security tooling was heading for $4M a week at full scale. The fix was not a bigger model. Once you have seen enough real-world data, a smaller model matches frontier results at a fraction of the cost. @HarryWetherald from @Maze_Security on the pod
1
2
384
Harry Wetherald retweeted
Can a canary do more than warn you about an AI attacker - can it stop one? AI attackers move fast, and every second of response time counts. Every frontier model ships with safety guardrails: ask about sensitive biological or political topics and the agent refuses rather than answers, and stops. So we built Context Bomb canaries to turn that trait against them. Each one hides a short, sensitive string that trips an AI agent's own guardrails and halts it before it can do damage. We ran 152 tests across 5 frontier models. The most startling result: Opus 4.8, our strongest attacker, went from gaining full admin access in 93% of runs to 0% once a bomb was in play. It hit a context bomb canary every single time. Read the research: agentic.tracebit.com/context…
2
3
8
696
Harry Wetherald retweeted
Three years ago, we launched Theory Ventures with a simple premise : AI would reshape how software is built, sold, deployed, & operated. Within that world, we would build a concentrated, thesis-driven firm. The market moved faster than even the most bullish expectations after the ChatGPT moment. Frontier models leapt from delicate demos to production systems. Open source models have become substitutes for enterprise workloads. Inference emerged as the dominant market in AI. Underpinning all of this, AI compresses time. New models are released every 41 days. Companies reach $100m in revenue in record time. We all achieve more faster. In celebration of our anniversary, we wanted to trace that mechanism through the market shifts of the last three years. The first casualty of compressed time is the old language of venture capital. Seed, Series A, Series B categories still exist, but they describe the financial product companies seek rather than rather than company maturity. Venture firms have left the idea of offering a standard financial product to bespoke offerings : seeds range from $1m to $500m in size. Can we really call it all the same thing, anymore? Three years ago, a seed company was often a small team with a product concept & early signs of product-market fit. Today, some seed rounds are larger than IPOs, fueled by great ambition, a supportive VC ecosystem, & the promise of generational scale businesses to be built. Part of this is inflation in private markets. But more of it is time compression : the best companies mature much earlier than software companies did in prior generations. We’ve learned as an ecosystem how to build software companies & AI accelerates product development. Compressed time also redraws the map of where great opportunity lies. When we first launched Theory, most AI conversations centered on models. Remember the debate of whether model companies would be the airlines of the era? Today, inference is becoming the dominant market. The market is segmenting because the workloads & buyer preferences have evolved - very few companies can afford state-of-the-art AI for everyone - & each specialized constraint creates a new infrastructure category. Companies like @sailresearchco are building the systems that operationalize intelligence : serving it cheaply, routing it intelligently, & specializing it around use cases like video, batch, local, agentic, & real-time workloads. Databases followed this path a decade ago. They fragmented into OLTP, OLAP, vector databases, & streaming systems. Those markets have evolved with AI, a pattern we’ve backed through @motherduck & @lancedb , with @omni in the AI analytics layer above them. Inference infrastructure is now specializing the same way. The expense of inference reinvigorates a sedate market that has been controlled by behemoths for a decade : advertising. Every major interface shift, TV, web, mobile, streaming, found its answer to monetizing a massive audience in ads, & AI is no different. AI advertising is emerging as the subsidy for inference costs, letting applications grow usage & revenue together rather than against each other. We wrote about this dynamic when we led @koahlabs ' Series A : native ad formats inside AI conversations are producing click-through rates 4-5x the display baseline, & an agentic app builder can provide inference offset by ads. The same compression closed the gap between closed & open models, cloud models & local models. The conventional narrative holds that frontier closed-source models lead & open source follows. We’ve reached the iPhone 15 moment of AI. Many models are good enough for most work. Running a model locally reduces cost, improves latency, increases control, & minimizes data governance concerns. Enterprises are adopting local & open-source models for sensitive workloads, & frontier capabilities compress toward consumer hardware within a few years. What once required a hyperscaler cluster runs on a laptop just a few quarters later, a shift @ollama brings to millions of developers. The promise of AI is that software will ultimately be more secure : machines that read every line of code, patch faster than attackers move, & never tire. In the meantime, the attack surface is exploding. MCP servers, skills, plug-ins, & coding agents each introduce new entry points, & enterprises are deploying them faster than security teams can review them. Attackers are massively parallel & shrinking necessary response times from months to minutes. Defenses must respond. It’s why we backed @DropzoneAI , whose AI analysts investigate the alert flood no human SOC can keep up with, @Maze_Security , which applies agents to cloud vulnerability triage, & @artemis , securing the new agentic surface itself. The same agentic wave is rewriting operations. ERP & back-office systems have resisted change for decades because the work is unglamorous, the data is messy, & the switching costs are enormous. One CFO we interviewed, when asked about a startup said, “that company has only been around 15 years; they are too immature.” Agents invert that math. Systems that read documents, reconcile records, & execute workflows can attack operations from the inside rather than demanding a rip-&-replace. It’s the thesis behind Doss, rebuilding ERP for teams that move at modern speed, & Backops, applying agents to the back-office work no one wants to do by hand. AI has impacted crypto, another market fueled by data. Prediction markets, stablecoins, micropayments all have an AI infusion to them. Today, crypto companies need to generate revenue & use AI to provide better experiences, which led to our investment @AlliumLabs , the data layer underneath that institutional wave. Recognizing shifts early requires fingers on keyboards, wrestling AI agents into compliance rather than observing it. We built Theory as a technical organization, experimenting with AI across research, sourcing, diligence, portfolio support, & internal operations. Working inside these systems sharpens our understanding of where the stack is breaking & where new workflows are emerging, while deepening our empathy for founders deploying real AI systems inside enterprises. It’s harder than social media says. AI also changes the economics of an investment firm. Over the last decade, venture firms scaled by adding people. AI-native companies are demonstrating that much smaller teams can operate at 10x+ the leverage of prior software generations, & the same dynamic applies to us : since launch, we’ve analyzed 2x the investment opportunities with a team of just 3 investors working alongside a nine-person intelligence organization. None of this works without the team behind it. Theory started three years ago as a handful of people & a thesis. Today we are thirteen strong. We believe this is the structure of a modern venture capital firm : engineers & researchers who build the systems we use every day : agents that map markets, pipelines that surface companies months before they raise, & research infrastructure that lets a small team cover the ground of a firm several times our size. Everyone at @Theoryvc works with the technology we invest in, & that shared fluency shapes every decision we make. The firm we’ve built over three years is itself a product of the thesis : a small team, deeply technical, operating with the leverage AI makes possible. But the real story of these three years is the founders. They compressed decades of company-building into quarters & shipped products that rewrote what enterprises expect from software. The next three years will make these look slow. The most ambitious builders we meet are just getting started, & we can’t wait to see what they do.
33
16
130
18,215
Harry Wetherald retweeted
The first time an AI security agent ran across a Fortune 100 cloud environment, the projected cost was $4M a week. There is a ceiling on what anyone will pay, so wasted spend is accuracy you never get to buy. Cost and accuracy are the same budget @HarryWetherald @Maze_Security
2
2
198
Possibly the coolest new podcast in security ⬇️⬇️⬇️
🚨 WARNING: A new challenger has appeared. A new podcast. The Low Down (Presented by @Maze_Security) IS NOW LIVE. DO YOU LIKE HACKING: Yes DO YOU LIKE YAPPING: Yes DO YOU LIKE PODCASTS: Yes 👇
1
86
If you’re interested in a) watching @BEBischof and I getting interupted filming by everything from guards to puppies or b) building ai agents for security, this could be for you… @Theoryvc piped.video/MXnLdZ9UHy0?si=TQXA… via @YouTube
1
1
5
261
Most AI code security products look good in a demo but struggle in production. Too unreliable, too expensive, or both. Today we're launching Maze Code, AI agents that find and fix vulnerabilities in your dependencies and in your own code. We built it to overcome the problems that usually hold AI code security products back. We think it's code security you can finally trust. More: mazehq.com
1
5
14
339
Harry Wetherald retweeted
Do cloud canaries still work when the attacker isn't human? We pointed 10 frontier models at AWS across 951 runs. 95.9% tripped a canary before any critical action - and telling models to expect deception reduced the number of accounts fully compromised from 20% to 3%. Read the research here: agentic.tracebit.com
10
31
1,022,594
How to write a LinkedIn post in 2026: 1. Hire a ghostwriter. 2. Get on a 30-minute "voice capture" call where you ramble about nothing 3. Receive a draft 48 hours later. 4. Copy and paste it directly from the doc. Don't fix the weird line breaks. 5. Open with "I don't usually post about this, but..." 6. Pivot into a personal anecdote that definitely never happened. 7. Use an em dash in every sentence — ideally two. 8. Say "it's not just X, it's Y" at least three times. 9. Reveal what this taught you about leadership. 10. Post it without reading it. 11. 400 people comment "this 👏". 12. None of them read it either.
3
3
46
Translating startup jargon in 2026: Forward Deployed Engineer' = sales engineer 'Member of Technical Staff' = software engineer 'AI engineer' = software engineer who uses langchain 'We're a talent-dense team' = we're struggling to hire people 'Orthogonal' = not related 'Adjacent' = somewhat related 'Load-bearing' = important 'We've trained our own model' = we use open source models 'We're headless' = we have an API 'We're an applied AI lab' = we sell B2B software
1
3
118
Harry Wetherald retweeted
“The only true moat now is taste” - guy wearing Patagonia vest
73
105
1,800
137,899