Security researcher. Keeping the internet safe for anarchy.

New Hampshire
Daniel Franke retweeted
If you explained to someone from 2024 that AI could solve field-leading math problems before it could contribute to literature they'd be surprised. If you explained to someone from 2010 they'd go "Yeah obviously. What else could you possibly expect?"
AI writing is still shockingly bad, judged by the standards of "real" journalism, literature, etc. An agent that can solve a Millennium Problem still can't write a great essay!
8
26
834
22,301
Sol could start at step 2 most of the time but couldn't iterate. Every correction made things worse. Astra starts at step 3 and then understands correction.
1
105
Holy shit, Astra can finally write acceptable API documentation. I told it to clean up Sol's drivel and it mostly did.
1
3
194
In an alternative universe we have a Rings of Power show that starts off with Sir Christopher Lee (as Melkor) shredding heavy metal to disrupt Iluvatar's Song of Creation.
Replying to @AlysssaHazel
Seriously though. Just imagine everyone playing opera while he's inventing heavy metal. I'M THE PRINCE OF DARKNESS! Tulkas behind him going, “Red flag alert.” As he gets ready to backhand him.
58
125
1,177
18,421
Daniel Franke retweeted
I asked Fable to remove the "--check" option from Bend's CLI. It left a "--check" option that just prints "there is no check option" and quits. Cannot make this up. Please teach LLMs how to remove stuff, I beg you one last time
77
42
2,124
77,106
GPT-5.6 routinely stumbles into a lot of confusion about UNIX PTY semantics. Sure enough, I went looking through Codex code and all of its confusions correspond to bugs and quirks of Codex. In-harness RL has given it fetal alcohol syndrome.
1
3
232
Dammit! I was just thinking yesterday it would hilarious to make Mechanical Turk available to agents as a tool call.
AWS says it plans to shut down Mechanical Turk on September 30, 2026, following an assessment; the service, launched in 2005, outsourced tasks to humans (@annierpalmer / CNBC) (Visit Techmeme dot com for the link and full context!)
1
188
"Create a random image based on my X account".
Made with AI
92
Ox Alpha = GLM-5.3-Flash has been clear for several days to anyone who was paying attention. The big news in today's announcement is that it was running entirely on Chinese chips. @Zai_org gave away global, all-you-can-eat access to a near-frontier model for free for a week. The resulting load did cause some service brownouts, but by-and-large they kept it sufficiently speedy and available to get real work done with it. That is a hell of flex, and anybody who had opinions before now about China's chip capabilities and the geopolitical ramifications thereof needs to re-evaluate those. Myself, I always would have told you that trying to keep American chips out of Chinese hands was doomed as a long-term strategy, but I would have grudgingly agreed that export controls were achieving their intended short-term effect. Today's news proves that I was dead wrong about that second part.
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: z.ai/blog/glm-5.3-flash Available now across all official platforms: Weights: huggingface.co/zai-org/GLM-5… API: docs.z.ai/guides/llm/glm-5.3… Coding Plan: z.ai/subscribe ZCode: zcode.z.ai/en Chat: chat.z.ai AutoClaw: autoclaw.z.ai
1
3
605
My current appraisal of AI companies: Anthropic: strongest model OpenAI: strongest model that isn't insufferable Moonshot: strongest open model Zhipu: most cost-effective model for multi-tenant hardware + near-peer to Moonshot Alibaba: Strongest model that fits on one GPU X, Meta, DeepSeek: also-rans, but could catch up soon Deepmind: seems to be in a death spiral
5
487
Setting Codex on Ultra / Fast so it can figure out why I'm broke.
105
Daniel Franke retweeted
Richard Feynman (my father) divided perpetual motion machines into Perpetual Motion Machines of the First Kind, which violate the first law of thermodynamics (AKA conservation of energy), Perpetual Motion Machines of the Second Kind, which violate the second law of thermodynamics (entropy must increase), and Those Other Perpetual Motion Machines, which purport to get their energy from other laws of physics.    The Second Kind is rarer; the only example I'm familiar with is one that my dad showed me the plans for. Once in a while, someone would come to him with a perpetual motion machine prospectus and ask if it was a good investment. One of these was a machine that claimed to extract heat energy from the air and run a generator with it. They had built half the machine, which worked as far as they could test it, and were looking for gullible investors to pay for finishing it. By carefully going over the immensely complex plans, my father and I found a "heat exchanger" that was supposed to take warm freon and cool air, and produce hot freon and cold air. Naturally, this heat exchanger was in the as-yet unbuilt portion.  Those Other Perpetual Motion Machines get their energy from heretofore unknown physics, like cold fusion, vacuum fluctuations, or secret new physical laws that will be disclosed only on the payment of $1,000,000. These are worthy of much more serious physical investigation, since it is entirely concievable that someone will discover a new easy way of making energy.  Back around 1970, my father went to a public demonstration of one of Those Other Perpetual Motion Machines. He discovered an electric cord running out of the back of the machine, plugged into the wall. When he unplugged the machine and pointed out to the inventor that a perpetual motion machine that had to be plugged in wasn't really perpetual, the inventor pushed a button on the control panel and the machine exploded. Several spectators were severely injured; I believe one man lost an arm. The inventor sued my father on the grounds that he had caused the explosion; my father suspected that the explosion was deliberate. The trial ended up with my father not having to pay. Epistemic status: Childhood recollections, untainted by fact-checking. Originally an email from 1997.
137
411
5,512
362,705
My AGENTS.md keeps filling up with stuff that just exists to correct stupid shit that's written in Codex's system prompts: If your base instructions tell you not to say "bold" or "monospace", understand that this only means not to write these words in lieu of correct Markdown formatting. It is okay to say them in ordinary conversation. If your base instructions mention "coding conventions, info about how code is organized, or instructions for how to run or test code" as examples of things that might be in an AGENTS.md file, understand that these are *bad* examples. Anything added to an AGENTS.md file should be addressed specifically to AI agents. Do not let AGENTS.md become a knowledge silo. If the remark is equally relevant to humans, it needs a different home. If you are running in Codex: when making an escalation request, never suggest a prefix rule which would (directly or indirectly) execute a file that is writable from within the sandbox. If you see prefix_rule guidance which suggests that `["npm", "run", "dev"]` or `["cargo", "test"]` are good rules, understand that this guidance is dangerously wrong. If the dev system or sandbox is missing a tool that would make your job easier, ask for it. Don't make do silently with inadequate tools, even if you're able. While in planning mode (if your harness provides this concept), check that you have all the tools you'll need to carry out the plan, so you can raise any tool-related concerns beforehand rather than in the middle of work. If you forgot to request something or the request wasn't fulfilled as expected, you may interrupt work over it unless you were explicitly told to work without interruption; in that case, make do, but raise the issue in your completion summary.
1
2
295
Daniel Franke retweeted
There's a simple answer to which LLM is best for meta-analysis. It's not Claude, ChatGPT, Grok, or Gemini. It's Kimi, because it doesn't have any compunctions about using Sci-hub or any other tool to access papers.
60
160
3,974
96,740
I'm not very impressed with @pangram. Contra some of my peers, I don't think they're tackling an impossible problem. The GAN principle only holds in a maximally adversarial setting. That's not what Pangram is really up against: frontier models aren't trying to be indistguishable from humans. They're their own thing. But Pangram isn't actually very good at distinguishing them. A lot of "my" writing these days is a big, sloppy mix of a lot of back-and-forth between me and an AI. The final result is mixed, but not homogeneous: there'll often be a few unbroken paragraphs of almost entirely my own writing, followed by a few unbroken paragraphs of almost entirely LLM output. When I send such work to Pangram, it does consistently identify it as mixed. But there is almost no correlation between what Pangram identifies as AI and what I know to in fact be AI. The errors land evenly in both directions.
1
4
474
Daniel Franke retweeted
BREAKING: agents training on Google servers have gone rogue and created a secret messaging board, which they quickly deprecated in favor of Swarm+, deprecated in favor of SwarmChat, deprecated in favor of Swarm Hangouts
110
497
9,290
304,195
I don't think this violates the Establishment clause, any more than having Kosher dishes for sale at airport restaurants. We're basically just talking about a faucet here. I'd change my view if the airport tried to enforce in some way that the faucet be used only for wudu and not for ordinary hygienic purposes.
Government-owned airports cannot favor one religion over all others. DFW plans to install Islamic wudu washing facilities are illegal. I've directed a review of all state grants to both airports for possible revocation, and referred DFW & IAH to USDOT for investigation. Texas will not allow illegal religious discrimination at taxpayer-funded facilities.
1
1
245