Former Scale AI founder and newly appointed
@Meta Chief AI Officer
@alexandr_wang dropped a bombshell during his YC conversation with
@garrytan, and you could almost hear the tech leadership world go quiet.
He said Meta has already seen this internally:
Build the right agentic loop, give it an evaluation system and metrics that let it optimize itself, and a group of AI agents can complete more work than a team of 100 senior engineers.
And they do it “very easily.”
But the most interesting part wasn’t the 100-engineer comparison.
It was how simple the system underneath it actually is.
You’d expect some insanely complex, almost alien architecture powering a swarm like this.
Instead, Wang described it with a few almost comically basic words:
“Markdown files, cron jobs, goal, metrics, data.”
Once you strip away the hype, the implications for traditional software engineering become pretty clear:
1️⃣ It’s not that the models are magically smarter. The eval loop is doing the heavy lifting.
Traditional approach: humans write prompts, run the code, inspect the output, and hope nothing broke.
Meta’s approach: turn the business goal into something a machine can score automatically.
The agent submits its work. The system runs tests, calculates metrics, finds what’s wrong, and sends it back for another pass. Repeat until it passes.
Nobody has to babysit every step. The metric becomes the supervisor.
2️⃣ The real alpha is burning 1,000x more tokens inside the feedback loop.
A lot of people are still optimizing for the cost of a single AI call.
The frontier labs are playing a different game: spend 1,000x or even 1,000,000x more tokens in the background so agents can constantly review, rerun, challenge, and verify each other’s work until they reach a reliable business outcome.
Token cost is fixed. The payoff is a pipeline that keeps running.
3️⃣ Memory doesn’t need some fancy database.
Persistent memory can live in Markdown files.
Scheduling can be handled by the server’s built-in cron jobs, running overnight.
The simpler the scaffolding, the more robust the system can be. Less infrastructure also means fewer ways for context to fall apart.
This is a pretty brutal change in how technical organizations work.
The ceiling for a tech lead used to be partly about how many people they could manage, how many meetings they could sit through, and how many teams they could coordinate.
The leverage for the next generation of technical leaders may look very different:
Can you turn a messy business objective into a rigorous set of metrics that an AI can evaluate automatically?
If 100 people’s output can be replaced by a few cron jobs, Markdown files, and a well-designed eval loop, the era of “just take the ticket and write the code” is coming to an end.
The people who can design the evals and orchestrate the swarm aren’t just holding a new tool.
They’re effectively running a virtual company.
Rohan Paul