Builds and operates billion-QPS distributed systems. Senior Director @MongoDB. Core infra, databases, reliability, and engineering management. Views are my own.

Dublin, Ireland
A special-purpose database optimized for a specific workload can indeed do better than a general multi-purpose one. The hard part is to operate this database with a similar reliability posture. Should review the cost of operating this and the uptime figures in a year's time.
We built a replacement for AWS DynamoDB, a key-value database for fast web content fetches. This was done with two engineers and hundreds of persistent Computer agents over two months. Migrating to our in-house database will save us up to a hundred million dollars yearly.
2
344
Highlights from the Aug 25 Buildkite incident: - Deployment caused an EKS pod to scale up. - Infra couldn't scale up in time. More scale up requests piled on. New pods put extra pressure on internal DNS. - Discovery slowed down by 780X and went down with a huge blast radius.
1
10
984
Today I started as a Senior Director of Engineering at @MongoDB. I'll own Observability Engineering. Ping me if you have pain points and/or ideas!
244
94
4,920
195,792
Today was my last day at Amazon after 6+ years. I worked on some of the world's largest scale distributed systems and learnt things I wouldn't be able to imagine before joining Amazon. I'll miss the daily thrill. Next week I'm taking on a new challenge and looking forward to it!
4
114
6,565
Highlights from the Aug 17 Github incident: * Load balancers got overloaded due to autoscaling config being insufficient. * Cascaded to more LBs, causing auth latency and failures. * Retries made it worse and recovery harder. So, yet another one for the Retry Storm Hall of Fame.
The RCA for the outage we had at GitHub yesterday is live. We're sorry it happened and hope this answers your questions about it. The team is already working on the actions listed here to fix it. githubstatus.com/incidents/z…
2
1
20
4,265
The retry storm was fixed by reducing gateway retry logic (with a PR), and blocking inbound requests at the load balancers.
6
1,922
The deeper problem is operational risks are probabilistic and the same risk may not materialise repeatedly, which makes it difficult to defend a big strategic investment to fix it. This turns the ops into an endless whack-a-mole via small tactical improvements that won’t hold up.
The problem with reliability work is that you can't tell the C-suite or board about how much revenue you brought in by stopping the system you built from dying, sometimes, but not 100% of the time.
1
5
2,734
Seniority is less about the code you write but more about the size of the outcome you're accountable for and how much ambiguity you can navigate while delivering it.
1
16
1,471
I wrote about how expectations change as engineers become more senior within an organization. Link below: eonem.substack.com/p/how-do-…
6
1,092
I follow the AI & math discourse a bit - without understanding the actual math. I thought I might offer some thoughts as a scientist in structural biology - a field transformed by various tech over the years, and recently by AlphaFold. It might perhaps interest @nasqret et al.🧵
I'm currently returning to Toronto from a summit on the future of mathematics, at OpenAI. @SebastienBubeck asked me to talk a bit about the future we'd all like to avoid, where humans are mathematically disempowered. @Jacob_Tsimerman advised us to try to prioritize detail over correctness, and I have no doubt that I succeeded in deprioritizing correctness. I tried to find a title that wasn't too bombastic:
10
69
341
70,667
You can in fact ask a candidate how they’d implement autoscaling for your database, then dive deeper into potential issues of the proposed approach. There are a ton of tradeoffs to discuss.
you can tell how inexperienced an engineer is by the degree in which they believe in autoscaling.
1
3
1,651
* How should it decide what scaling action will help the situation at hand, e.g. a hot key read scenario so create more read replicas, or a hot key write scenario so re-partitioning is required?
1
163
* How should it ensure the customer won’t get throttled or see degraded performance while the autoscaling is happening? * What other capabilities of the Control Plane should it take a dependency on? What happens if those capabilities go down?
1
73
Evren Önem retweeted
The elimination of engineers was one of the wilder hypotheses out there. Just like absurdly wrong. We just gave engineers a power tool that can accelerate the development of whatever we want. Of course their value not only remains critical in that world, but actually goes up in many domains because we can apply engineering to far more work than before. If you’re trying to automate drug discovery, you need engineers. If you’re trying to automate manufacturing, you need engineers. If you’re taking on larger and larger software projects, you need engineers. This will happen in domains beyond engineering too. AI causes companies to take on way more work than before, which leads to needing more experts to oversee the work. Even as models get better and better, they can be better utilized by the experts in those fields than the novices. Great time to be an expert.
Have you noticed that people have stopped saying that software engineering is over?
90
78
640
117,175
One take-away seems to be that if your deployment requires taking hosts down, the process should first check if the remaining capacity suffices to serve the traffic. githubstatus.com/incidents/q…
1
932
“My services are well and healthy despite not reviewing a single line of LLM-generated code”. Meanwhile, the production system:
Tarladaki çocukların çamur keyfi… (Şanlıurfa / Siverek / Karaca Dağ)
1
7
1,316
Code != value. You can ship 1M lines a day and still ship absolutely nothing anyone needs, or an unreliable product that breaks often. My favorite species is the “I don’t waste time reviewing anything, I use that time to ship more” bros. These guys ship tomorrow’s archaeology.
we're in a strange transitionary period where people are confusing what was historically scarce with what is actually valuable. for the last few decades, knowing how to code was highly correlated with making a lot of money and building valuable things. so now that AI has given everyone the ability to code, the obvious reaction has been to start building software. but the value of coding was partially downstream of the fact that most people couldn't do it. which makes vibe coding kind of paradoxical - millions of people are rushing to acquire the output of a skill at the exact moment that output is becoming abundant. i think the trend dies down once people realize that not knowing how to build software was never really their problem.
11
919
Copying & pasting from “the machine” or blindly doing what “the machine” says can not make you any more intelligent. Without a well-informed, insightful perspective you truly own, it’s harder than ever to distinguish yourself from the crowd.
We spent our entire lives, and human history, under the assumption intelligence was a scarce resource. Now, everyone can summon an inexhaustible IQ machine like it's tap water. And incidentally, the change is so profound our own brains can barely comprehend it.
3
624