If someone runs an eval that costs millions in compute, an agent takes a detour out of the sandbox, and now millions of dollars of compute are pointed at hacking you, I have bad news without LLMs, if someone was willing to spend millions targeting you, they were probably getting in. This has always been the dilemma with nation states like China. If someone can dramatically outspend you, well…
I realize there’s a separate argument around local models and agents, but I see leadership at all kinds of companies suddenly trying to defend against whatever swarm scenario matches the latest high profile incident.
You should probably worry more about someone pointing a 27B local model at your home built, public facing web apps than the first scenario.