Co-founder, CTO @buildkite, Dad, Llama Aficionado (not the model). Previously @GroqInc @PatientNotesApp @cashapp, @99designs.

Victoria, Australia
Lachlan Donald retweeted
Test workloads are build workloads; test bottlenecks are build bottlenecks. It's time for a new kind of scheduler, that can run all our builds and tests as an integrated workload rather than dislocated siloes. Then we'll wonder how we ever tolerated CI this slow and wasteful.
3
5
13
4,006
Lachlan Donald retweeted
42
467
6,364
215,628
Lachlan Donald retweeted
An OpenAI agent swarm operating under poorly airgapped research conditions accessed my CBA Mastercard and bought 4 schooners of Newtowner between 5:52pm and 7:15pm last night
77
549
10,929
239,027
I'm looking for someone brilliant to come run internal developer velocity (DevEx) at @buildkite with me. Come help us build faster developer loops and ship software for the frontier labs and more. If you love CI/CD and agentic software development and have a track record of high-agency dev tooling at scale organizations, look me up for a chat. job-boards.greenhouse.io/bui…
10
16
103
42,489
Lachlan Donald retweeted
Accidentally said CI/CD instead of software factory and they kicked me out of SF
accidentally said microvm instead of agent sandbox and they kicked me out of SF
52
183
3,380
165,256
There's a lot of talk about software factories right now. Frankly, most of it is bullshit. Beware “I do AI better than you" snake oil and think more from first principles. Strip most of them back, and you find a series of automations, and one of them happens to be a coding agent. (Which is also what CI/CD was. You stopped clicking around testing by hand, and a machine did it.) There’s a lot right with making your idea → ship loop continually faster and smarter, but I’m not sure that’s really a new idea. Nobody starts with a factory. A chocolate factory starts as one person and some beans in a kitchen, hand-making chocolate. Then they get a mixer. A bigger fridge. A conveyor belt. Then a packaging machine. Each addition builds scale, consistency and efficiency. A small business becomes a medium business, which becomes a factory. What changes along the way? Where the person stands. What the person does and the technology they can access. First, they do every step. They buy a machine or two. Then they do some steps and watch the machines. Then they operate and connect the machines. Then they redesign the factory itself. Software is the same loop. Ideas are dreamed up, feedback and operational bugs come in, someone turns them into a list of work items, someone builds it, someone reviews it, someone builds it, someone deploys it, it ships, something breaks, someone notices, someone files the bug. Rinse. Repeat. You've run that for years with a person at every station. To improve the factory you already have, upgrade one station today. Add automations and AI agents to code or review. Observe, learn. Then do the next one. Thoughtful, ambitious, patient scale. Are there roles for humans in that factory? Abso-bloody-lutely. Initiating ideas, observing the operation, reviewing at discrete checkpoints or severities, co-designing both the flowed work and the factory itself, owning exceptions and resolving conflicts, adding context. (Just look at one of our latest Loom action plans demo👇🏻 for an example of human-AI collaboration) I believe this human judgement and intuition only get more important as factory throughput scales. In short - don’t believe in the mythical agent month. You already have a software factory. AI just allows you to make more steps repeatable, more flexible - and that lets you move where you stand.
Replying to @SpaceXAI
Atlassian used Transcribe 2.0 in @Loom and found it more accurate at capturing user instructions. Dictate your change requests using Loom, then directly export to Cursor to code it up.
27
17
261
42,524
Test selection is going to be huge for compute reduction
Replying to @ggsimm @alxfazio
We already run this at @buildkite — same approach as meta, xgboost ranking. −53% of tests on feature branches. testing jev against it now on recall and latency. not sure how we get past it collapsing into a path/diff similarity heuristic though…
1
8
1,122
Lachlan Donald retweeted
🥳 Excited to start revealing what we've been working on in the last few months. First, we decided to reinvent Kubernetes for agentic workloads with statefulness and fast resumption. Secondly, we are building an agentic orchestrator that will be Google's open agentic orchestrator and runtime. github.com/google/ax
128
378
4,237
717,565
Mostly agree on these. My only reservation is on unit tests vs more complex integration tests and how that works as compute becomes the shortage.
Most predictions I see are still way too conservative. Here's mine
3
7
1,694
Lachlan Donald retweeted
new post: the senior engineer death spiral sunilpai.dev/posts/the-senio… a friend just started a big job and asked for some advice. so I braindumped a monologue about a super common failure mode I see with engineers and posted it here, hope it helps whoever it can.
280
640
7,258
1,825,789
Lachlan Donald retweeted
74
661
8,600
253,861
Lachlan Donald retweeted
I think I speak for the internet when I say: fucking finally
We're adding support for AGENTS.md to Claude Code. Starting today in version 2.1.277, if there is no CLAUDE.md in a folder, Claude will check for and use AGENTS.md. You can toggle this behavior in /config.
5
3
68
5,476
Agreed, this is my expectation too. Architecture, trade offs and assumptions.
The "whiteboard defense:" I should be able to pull you aside at any moment and ask you to explain any customer-facing system you've shipped. You should be able to clearly explain how it works and defend the decisions you made. This is my benchmark for responsible AI usage. I don't expect line-level familiarity with the code. I don't care if you remember the exact function name or implementation detail. You may not even know it. I don't care. But if I ask "why did you do X instead of Y?", "what happens if this actor behaves maliciously?", "what data structure did you use here and why?", or "where does this fail?" you should be able to answer confidently. For PoCs, demos, experiments, whatever: I don't care. Generate 100% of it and understand none of it. Speed over quality every time in those specific scenarios. But if you're shipping customer-facing work, you can't be shipping things you don't understand at a high level.
1
1
10
586
Lachlan Donald retweeted
A very interesting second-order affect of our PR approval agent is that the feedback from the agent given before approvals has caused reverts to dramatically drop. I suspected that our code quality has improved a lot from my own use of it.
Prepping slides for my talk at SignalsConf – here's an interesting one – our PR revert rate has dropped 3x right after we rolled out our latest generation of the code review and approval agent!
1
2
20
1,713
Lachlan Donald retweeted
I can’t stress enough how little an idea matters compared to the agency of the people executing the idea. I have had the privilege of knowing and sometimes even working with some of the most successful people (by various metrics). The difference between mediocre and excellent work and outcomes is predominantly one of agency. In practice this means: they dont wait for things to happen to them they go out and make things happen for them. They don’t wait for someone else to do something, for someone to teach them, for someone to give them the path, etc. They just go out and find a way to do it. I think the single biggest superpower these people have is the realization/belief that the world around them is completely mutable. Most everything that happens is because a person made it happen. I used to tell people to look around the room you’re sitting in. Look at everything. Every noun. It almost all exists because a person willed it into existence. Nothing is stopping you from doing the same. I see people online all the time dismissing someone else’s success because “I had that idea first” or whatever. I mean… yeah? If so then the difference is… you. So a bit of a self own whenever I hear that. Number one tip: act with agency.
228
1,279
11,856
682,069
Hey @pangram I'm not sure if it's a humblebrag, but requiring a sales call for org subscriptions and then not having availability till November isn't great. Any other ways to send you folks money?
9
426
Lachlan Donald retweeted
The benchmark results are telling us what we sort of already knew: patching vulnerabilities without breaking something else can be trickier than identifying and exploiting them. That's why it's so much better to find any vulnerabilities as early as possible in the software development cycle. What's interesting to me in the age of LLMs, however, is how much easier it is for LLMs to automatically discover and exploit vulnerabilities than patch them correctly. Vulnerability exploitation can be evaluated as a closed-loop problem, but most software doesn't have pervasive enough tests or specification to make patching the vulnerability a closed-loop problem. LLMs can only guess at open-loop problems based on their training.
🦔AI can only fix security vulnerabilities 26% of the time. Researchers at 1Password ran over 6,000 AI-generated patches using Claude and ChatGPT against real vulnerabilities. Half the time the AI failed to fix the original bug. 4.5% of the time it created a brand new vulnerability that didn't exist before. And the researchers found that checking an AI-generated security patch takes more effort than just writing the fix yourself. Their conclusion was blunt. "The expected value of a fully LLM-generated, non-human-reviewed patch is a net-negative by a considerable margin." My Take OpenAI launched a program this summer called "Patch the Planet" where AI finds bugs and generates the fixes. These researchers ran 270 patches against one of those same bugs. Zero clean fixes. Not one. Every patch that fixed the original problem created a new vulnerability in the process. The partner that submitted a fix through OpenAI's program produced what the researchers classified as the worst possible outcome, it didn't fully fix the bug and it introduced a new exploit on top of it. Here's what this means if you don't write code for a living. Companies are using these AI tools to patch the software that runs your bank, your hospital, your phone. The pitch has been "AI finds and fixes security holes faster than humans." This study says the AI fix is four times more likely to be broken than working, and a third of the time it recreates the exact same mistakes human programmers already made. Reviewing the AI's work takes longer than doing it yourself. So the speed advantage disappears the moment you try to verify the output, which most companies won't do because the entire point was to move faster. I think we're going to see major breaches traced back to AI-generated patches that nobody checked, and the companies that shipped them are going to blame the tool instead of the decision to trust it. Hedgie🤗 Study: 1password.com/files/resource…
7
9
42
5,484
Wow TIL
穴に落ちるキノコは、上からとれる!
7
645
Lachlan Donald retweeted
Beware people obsessed with outcomes instead of building outcome machines. Its worse than ever with AI, but these people existed before. Short term results above all else, etc. Don't fall into the trap. Invest in building strong fundamentals, invisible supports, and outcomes flow like water. An outcome machine.
115
552
6,340
271,881
Lachlan Donald retweeted
Switching between threads in Amp is now a lot faster on Safari & our macOS/iOS apps Using @buildkite preflights (buildkite.com/docs/platform/…) on a macOS runner against the orb's tunneled dev server, opening Safari and collecting a web perf profile. Found lots of CSS/layout speedups.
3
8
63
3,815