🦘/acc + CTO of @HelmGuard. I sometimes post about deep learning.

London, England
There's a theory of AI adoption that it will come in two waves. In the first wave (phase 1), businesses take an existing process, and automate parts of it with AI. In the second wave (phase 2), people will replace the existing process with something adapted to the new paradigm. What will this look like in cybersecurity? Do we have a sense of what a phase 2 use case might be? I think a very good candidate is third party compliance and security. Today AI is being used in the space to speed up an existing system: you design an agent to be 10x faster and cheaper at extracting claims from a SOC2 report or reviewing the outputs of a penetration test. I think there is an entirely new system that can be built around the idea of "stateless auditing" or as we call it at @HelmGuard ephemeral verification. This leverages the fact that AIs can forget the information they see. They can verify facts about data without revealing that data to the verifier. I think that in the future, if you want to verify whether a supplier you are working with is secure to your standard, you won't ask them for a SOC2 report. You'll send a specification of what you consider secure to a trust agent in their environment which will verify those claims for you and return a yes or a no.
1
56
One underrated aspect of critical infrastructure security is the susceptibility of these orgnaisations to supply chain attacks. If I were a threat actor, I wouldn't look to compromise a public healthcare provider or energy grid operator directly, I'd infiltrate their supply chain. What makes this easier than in most other sectors: 1. Critical national infrastructure often has a public record of all suppliers used (this is usually to keep procurement fair) 2. Suppliers are typically legacy companies with very poor security 3. The number of suppliers is enormous Instead of going after infrastructure where there have been heavy investment in internal security, I'd go after the supplier they've been relying upon for 20 years that hasn't gone through a security due diligence process.
1
60
As we contemplate a slow-down on model progress. I think it's important to ask what a pause would mean for cyber defence. As the co-founder of a cyber defence company, I personally think there's years of progress we could make in elicitation and diffusion if we froze at Fable-level capabilities. For example, looking at one type of work security teams need to do on the regular: evidence collection and control mapping. Organisations have a bank of controls that they use to regulate what is permissible e.g. "all privileged accounts must have MFA" or "all development work must be done on a remote VM with a secure OS". Practitioners then need to manually assess whether this control is enforced across all of the relevant assets an organisation owns or across all of the people that work there. Today there are too many assets, people and programs for these teams to assess, so they sample a controls' effectiveness or use deterministic rules. However, this work is relatively easy for an AI model, it involves some multi-modal understanding but primarily reasoning across textual inputs. On our evals, this sort of work is easily handled by the Opus 5 and Terra 5.6 class of models. It would be a phase change for cyber defence if controls were assessed at a population level not by sampling. That is by looking at everything an organisation is doing. This is just one example, but there are countless others where a pause wouldn't impact progress.
1
3
133
Jack Miller retweeted
Is GLM-5.3-Flash Mythos-level at cyber? In our latest post, we find that GLM-5.3-Flash can outperform Mythos Preview on ExploitBench at ~6% of the price. We do this by scaling up inference tokens to 1B per vulnerability.
1
2
5
129
We’ve just raised! HelmGuard be growing and protecting
We've just raised $7.3M in a round co-led by Infinity Ventures and Frontline, with participation from FinTech Collective, Stage 2 Capital, and Entrepreneurs First. Read more here: helmguard.ai/resources/7.3m-…
1
9
258
Jack Miller retweeted
I’m increasingly worried about misalignment that subverts training itself. Following our ICML work on exploration hacking, we study what makes training setups vulnerable, starting with AI debate. Two posts w/ Jason Brown, @BraunJoschka, Roland Zimmermann & @davlindner Links 👇
At ICML, I presented our Exploration Hacking paper with @DamonFalck and @n_lie_k. Thanks for the discussions! Two questions stood out: 1. What makes exploration hacking more likely during RL post-training? 2. How else might models strategically influence their own training?
4
2
22
826
Jack Miller retweeted
If you believed what Anthropic leadership believe, how would you have released Mythos? And what would you say publicly + privately about AI risk?
1
10
991
Jack Miller retweeted
GPT-4.5 remains the most criminally underdiscussed launch in the modern history of AI, relative to its significance. 4.5 was the death of the second, Secret scaling law - the hitherto persistent correlation between decreasing cross-entropy loss on next-token prediction and "usefulness to the consumer". I see its failure as the definitive pivot away from attempts to one-shot superintelligence out of the box towards a more modular conquest of the intelligence tree (through task-specific RL).
3
7
43
3,083
I’m quite into this idea
Wacky but secretly-good idea that follows from taking the whole "AI constitutions" premise seriously: There should be "courts" to make rulings on how the Constitution/Spec applies to contentious instances of model behaviour. Users should be able to report + contest costly refusals through a button in the app - most of these are ofc auto-resolved with AI, but the interesting edge cases bubble up through a hierarchy of courts, with the most contentious of all being resolved by a "supreme court". Many advantages to this approach: 1. Claude's ability to apply the constitution to a given situation will be massively, massively improved by having a rich body of precedent to contemplate and refer to. Such a body of precedent makes the constitution a lot "thicker" - it could also be open-sourced to improve the (currently somewhat disappointing) transparency into the Constitution, allowing users to know what to expect from Claude. I imagine that Amanda, Joe et al currently produce such "precedent" for Claude with synthetic data and galaxybrain theorycrafting, but I'd happily bet that the real world is a better, richer source of edge cases than anything even they could come up with. 2. This could be a really, *really* interesting way to get the "democratic input" into AI constitutions that everyone keeps clamouring for. Usually these proposals end up as uninspired calls for using surveys and focus groups in the drafting stage of the Constitution, which I think is a fairly limited way to think about the strengths of liberal democracy. On this "courts for the spec" proposal, you could imagine opensourcing stages of the judicial process in wide variety of ways. One thing you could do is to crowdsource amicus briefs (or perhaps even the whole case!) for petitioner or respondent. I feel like there's a promising Pettit angle to this. 3. Discursive and argumentative traditions (and not just surveys or isolated technocrat drafting) have a good track record for being the means by which we as humans resolve these kinds of problems, so it just makes sense to get this going for AI. I particularly think that debates are likely to be much richer when they are about *specific* instances of model refusals than just vague, open-ended discussions about how AIs should be governed. It's also predisposed to iteratively co-evolve with the pace of technological change far better than any one-and-done philosophising would 4. There may even be some kind of "separation of powers" argument to make here, insofar as the "legislature" drafting the spec/constitution is distinct from the "court" ruling on how it applies to a particular case. More applicable to OAI than Anthropic (which is more generally comfortable with moving past traditional liberal principles) ofc. 5. Finally, and most obviously, there *should* be a way to contest model refusals!!! If we all take seriously the idea that agents will become some large % of the human economy, then an individual unjust refusal could be absurdly costly. I think there should be a transparent, participatory way to contest this such that we are not at the mercy of the AI conglomerates. I cannot think of a better way than to borrow from the law. Perhaps I'm being too Anglobrained with this, but I just think the law and courts pareto-mog almost all other forms of value-resolution we homo sapiens have ever come up with
153
Jack Miller retweeted
Most of us know by now that OpenAI and Anthropic take different approaches to governing the AIs they create. In crude terms, OAI constrains GPTs with rules while Anthropic steers Claude with character. Which approach is preferable? My previous attempts to consider the question have indulged in (aspirational) nuance. Today, I'll attempt to explain the trade-offs involved with the crudest possible heuristic: Would you rather risk a Robespierre for a chance at Petrov, or Pontius Pilate for a shot at John Adams?
1
4
36
3,705
Jack Miller retweeted
To recap: you can imagine two ways to create AIs that are faithful servants of mankind 1. List all the things that the AI should not do, and tell the AI not to do them. or 2. Make your AI have a personality so dope that you trust it not to do bad things OAI opts for 1, Antropic for 2. Anthropic's Joe Carlsmith is appropriately grand about the scale and significance of this decision
1
1
5
646
Jack Miller retweeted
How should we think about the advantages of trusting character over rules? Trusting an individual's judgement has a lot of clear perks. Rules, like Fallen man, are inescapably imperfect. Strict rule-following can lead to perverse, negative-sum outcomes intended by no one. There's a simple workaround - just rely on the discretion and conviction of the righteous! Consider the almighty Stanislav Petrov (to whom we owe so much, if not everything). His rule-breaking is all that stood between our civilisation and a cold, empty future as stardust.
2
1
5
571
Jack Miller retweeted
And yet - the history of human discretion is not populated with Petrovs alone. To rely on the soundness of an individual's discretion is to fatally put oneself at the mercy of the mad, reckless and cruel. It was discretion, license and conviction that led Robespierre to the Terror, to the massacre and bloodthirst of the Parisian summer of 1794.
1
1
3
477
Jack Miller retweeted
And what about rule-following? Do we not owe a debt to the law-abiding pedants of history? Perhaps because our global (Californified) culture worships mavericks more than company men, the heroic sticklers of history go by largely uncelebrated in their unvisited tombs. As a counterpoint, consider the upright fastidiousness of John Adams (your favourite founding father's favourite founding father). When British soldiers fired upon the innocent crowd at the Boston massacre, Adams - to the furious displeasure of his countrymen - abided by his duty as a lawyer, took up the soldiers' legal defence, and won acquittals for six of the eight. He did so *in spite of* the moral instincts and revulsion which told him the soldiers must be executed, deferring instead to the tired letter of the law. Adams would later call it one of the best pieces of service he ever rendered his country.
1
1
9
658
Jack Miller retweeted
But the sticklers of the earth also shoulder the blame for history's most heinous crimes. To illustrate this point, we need look no further than the Gospel of John. Take the actions of Pontius Pilate, that ill-famed Prefect of Judaea - slavishly sending the Son of Man to his death at the hands of a mob he knew to be wrong in their condemnation of the itinerant Nazarene. The murder of Jesus is a fairly costly failure mode.
1
1
3
555
Jack Miller retweeted
Now we can return to the respective approaches of our contempory billion-dollar aspiring Demiurges. OpenAI shoot for the reliable rectitude of Adams while paying the price of a potential Pilate. For Anthropic, the chance of a Petrov is worth the risk of a Claudespierre.
1
5
21
1,794
Jack Miller retweeted
It's increasingly hard to distinguish genuine alignment from alignment faking and reward hacking. Behavioral data isn't enough, so we need to map models' motivational structures! What latent structures track model motivations? @DavidDAfrica and I lay out this problem.
5
12
101
5,675
Jack Miller retweeted
Incredibly fitting that Grok, despite being pretty mediocre at forecasting overall, is the single most useful addition to an ensemble of frontier models forecasting real world events. Overwhelming empirical evidence in favour of keeping your toxic friend in the group chat
15
26
1,058
78,877
Jack Miller retweeted
OpenAI's Model Spec contains a set of principles they commit to upholding across all deployments. These include pledges against enabling acts of violence, the creation of WMDs, and use of their models for "mass surveillance". What is the nature of these commitments, and what mechanisms exist for upholding them? #5 in my series on the Model Spec covers OpenAIs "Red-line principles"
2
3
13
813
Jack Miller retweeted
More on OAI's model spec. This time, I'll cover the section that most distinguishes OAI's liberalism from Anthropic's Claude-as-philosopher-king approach. #2 - "No other objectives" Almost infinite material for discussion here
14
28
296
66,527