ML systems & agents infra @linear, fmr @meta. @uwaterloo SE alum

nyc/toronto
Pinned Tweet
I post-trained Qwen3-Coder to fix bugs using an actual debugger. The result: Solve rate: 70% → 89% Median turns to fix: 46 → 19 (-59%) Instead of just reading code or print-debugging, it: - reasons from execution - inspects live variables and call stacks - sets breakpoints, steps, and evaluates expressions
91
117
1,566
127,999
mufeez retweeted
It passes the test! Send this post to your agent and watch it work wonders on your CI times – also read and learn ;)
Agents were shipping code faster than our CI pipeline could keep up. Our team optimized our pipeline, leading to a roughly 15% faster test suite and 50% reduction in runner-time spent per test, all while our test suite grew 4x in size. @moofeez explains how: linear.app/now/ci-bottleneck…
2
3
152
70,015
Our test suite is growing 9x faster than a year ago, now that agents write most of it. If we'd left Cl alone, a PR would wait ~11 minutes today. We got it to 5, on 3.7x the tests. Give the blog post to your agents for free gains!
Agents were shipping code faster than our CI pipeline could keep up. Our team optimized our pipeline, leading to a roughly 15% faster test suite and 50% reduction in runner-time spent per test, all while our test suite grew 4x in size. @moofeez explains how: linear.app/now/ci-bottleneck…
10
3
64
12,250
mufeez retweeted
The test of a great technical blog post at this point is: can I point an agent at it and have the learnings translated across to my codebase?
2
2
17
1,853
really not liking the trend of code review becoming agent ping pong…
22
9
299
45,316
I do not subscribe to this mode of operation
5
3,271
btw make friends with your AWS/GCP rep now if you want compute next year
18
1,111
how long until my Instinct can talk to your Instinct directly
2
1
10
1,952
imo Instinct worth at least 5B after they implement this
2
447
Recently crossed 3 years at Linear Wild how far a cracked team + beloved product takes you 🫡 (we're hiring!)
Today we're announcing our second employee tender offer. We passed $100m ARR earlier this year and now have more than 40,000 paying customers. The tender lets our teammates participate in that success at a $2.5B valuation: linear.app/now/sharing-growt…
2
73
3,186
Kimi K3’s RL setup is surprisingly “boring” synchronous RL, specialist policies, MOPD maybe a great pretrain is all you need
Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. New model architecture: 2.5x the intelligence per unit of compute, not just more params. Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale. Model weights: huggingface.co/moonshotai/Ki… Tech report: github.com/MoonshotAI/Kimi-K… Tech blog: kimi.com/blog/kimi-k3
6
927
ending the week with a ~36% drop in median CI time 🫡
1
9
696
Proud to say that Devin has cracked three more unsolved problems today 1) REFUTED: Graffiti Conjecture 154 (open for ~40 years) 2) PROVED: Graffiti Conjectures 39 & 40 (~40 years) 3) REFUTED: Brandt's Regular Supergraph Problem from West's open problems list (~20 years) My methodology is explained below, but basically I showed Devin the original tweet and told it to find similar problems and crack them. The flood gates are open.
Community note
Graffiti Conjecture 154 was already refuted on June 11, 2026, prior to this post claiming it unsolved. github.com/demonstrandum-…
1
12
1,348
spent some time this weekend implementing IPPO and MAPPO and trained two agents to run an Overcooked kitchen together fun watching them slowly learn to cook 🧑‍🍳
1
2
15
890
the same setup learned a different coordination strategy across all three layouts: - cramped_room: avoid blocking each other - coordination_ring: move in the same direction - counter_circuit: coordinate around the loop
2
199
would recommend making a lil rollout inspector for debugging RL trajectories 👍
5
548
Replying to @nico_laqua
I get that business insurance is similar Nobel level type of pursuit as ground breaking physics and the Manhattan project. Hopefully the blast radius will be contained. I don’t think the disagreement is whether hard problems require intensity. The disagreement is whether intensity has to become a permanent operating model, and whether working seven days a week is the thing that compounds. My argument is that for most startups, the real compounding advantage is not raw hours. It is clearer thinking, better judgment, learning, and a team that can sustain high-quality work for a long time. You can always spend a lot of time working, but the PMF might never arrive. There are moments where extraordinary effort is necessary. Launches, incidents, existential deadlines, customer commitments. Those moments matter, and great teams rise to them. But if the company requires heroics every day of the eek, that usually points to a system problem. It means the operating model depends on burning reserve capacity instead of building it. Company that is constantly on fire is company that is not operating well. Whenever you put something out there, people will argue and people can argue the way I run Linear. The reason I comment on these things to offer some counter point. There is a growing cliché in startup culture where founders and startups feel the need to perform intensity publicly. How hard they work, how little they sleep, how many tokens they spend, how busy they are, how much personal sacrifice they make. You almost never see this from the most successful companies or people. Even if they work that way, they usually don’t make it the story, because they have more important things to talk about, like the product, the customers, the insight, the strategy, the quality of the work. That’s my issue with the narrative and why I think startups shouldn't blindly follow it. Not that is bad to work hard but grindmaxxing narrative can become the greater goal and become counterproductive. The performative intensity becomes the thing, and loosing sight of what actually matters. Lets check back in 7 years.
55
143
3,643
378,105
claude-squad is somehow still growing a year later <1 month of work btw
3
1
11
793
ok london kinda mogs nyc ngl
71
292
7,149
222,358