Performance Engineer. Tokyo based 🇯🇵

Tokyo
Eventually I imagine models will have enough intelligence and prowess to imbue good style, design and iterate until the simplest / most compressed solution is found (or humans are so out of the loop, that the latter doesn't matter). But I reckon in the intermediary; it's going to be very difficult for engineers to manage - offloading all work / thought to agents will result in total context loss, coupled with creeping complexity, slop & boilerplate: productivity may hit an asymptote soon and even decline. Incentive structures are currently as such, that this may be increasingly difficult to mitigate.
160
Oh you poor thing
324
1,696
94,278
4,946,224
Lazy work used to mean too little output. Now, with AI, it often means too much and more work for everyone else. @tobi call it "slop grenades." A "Slop Grenade" is when you let AI produce the work and pass it on without adding any value (including checking it). Someone else has to wade through it, catch the mistakes, and clean up the mess. You save time and look productive but someone else pays for it.
My third conversation with Shopify co-founder and CEO @tobi. 0:00 How Shopify Uses AI 7:18 River: Shopify's Internal AI 8:55 How to Encourage Osmosis Learning 10:52 AI Dreaming and Self-Reflection 11:53 How to Use AI for Strategic Decision Making 14:11 The One Thing AI Cannot Do 16:04 What AI is Making Worse at Shopify 19:46 Predictions: Where AI is Headed Next 21:55 The Future of AI-Powered Software 24:40 Will CEOs Be Replaced with AI? 27:54 Can Superintelligence Be Controlled? 31:13 Critical Skills in AI Age 34:22 Why Complex Solutions are Usually Wrong 36:33 Conditions Needed for True Intuition 38:02 The Best Path Doesn't Have Instant Feedback 44:51 How Affirmations Can Shift Your Behavior 50:44 The Inobvious Thing Hurting Companies 52:37 Relationship Between Beauty and Creation 56:23 How SpaceX Moves Forward By Subtraction 1:00:24 Why Companies Need Refounding Events 1:01:50 Books as Cheat Codes 1:02:39 Three Books to Change Your Thinking Enjoy! (Includes paid promotions.)
148
652
4,828
1,485,336
I've set up a rust framework for an attempt to get to ~cuBLAS perf for matmul
4
7
1,732
For the record, multiple simd registers + fused FMA gets us to ~15% of theoretical upper bound in terms of GFLOPs on the host machine for a single core. Still some way to go. Also, its already naturally following that tiling lends itself nicely for multiple cores.
1
64
Single tiling provides about a 30% increase in perf. Lower than I had from theoretical perspective. Tile sizes targets L1 cache size. Probably need to do something similar for SIMD registers & reduce some of the tiling bookkeeping - L1 cache misses reduce by 4x, yet instructions per FLOP increase by ~same amount, likely some inefficiencies, although when measured against control (TM,TN,TK=N) suggests bookkeeping cost less than 1%.
31
Astra has lost the plot
2
155
Ten years ago, I stopped using P99 and switched to P100, always. - it’s easier to calculate (just take the max) - it’s easier to understand (this is as bad as it gets, there’s nowhere to hide) - optimizing for P100 tends to optimize for all percentiles (especially over time)
Most developers know that avg latency is misleading and they should use percentiles instead. They also know that if 90tile response time is 200ms, then 10% of the requests had response times higher than 200ms. Well, the last sentence is false. Or rather, true in theory. The only way to calculate true percentiles is by sorting and scanning all the data. Since this often means many millions of datapoints, we actually don't do that. We approximate. With histograms. How does that work? You create buckets of values: 0 to 0.05, 0.05 to 0.25, 0.25 to 1, 1 to 5, 5 to infinity. Then you accumulate: number of data points, sum of values and the number of data points that fall in each bucket. Now, if I ask you for the 90tile value, you start going over the buckets from lowest to highest latencies, and add the number of points in each bucket. Once you go past 90% of the data points, you know that your 90tile value is in this bucket. Then you give either the middle of the bucket or the upper bound as the estimate. Note that you don't know what % of the data is higher than the estimate. And also note that the quality of the estimate really depends on how the data ends up distributing between the buckets. It isn't uncommon for ntiles 90, 95 and 99 to all fall in the same bucket (if your distribution is "heavy tail"). And sometimes all or most of the data falls in one bucket and you can tell nothing. I recently had a case where we noticed the 99tile of response time for a specific request type was 10s. This was both high and suspicious - we have a 10s timeout in some places. But we couldn't find any request that hit these timeouts in our logs. Super weird! Until a smarter colleague realized (by looking at sample traces) that 10s was just the lower bound of the top bucket. The requests actually timed out at 60s. This allowed me to find the timeout they actually hit and fix the issue. So the lesson is: Averages are misleading. But percentiles are usually approximate and can be misleading too! Be careful not to trust the data blindly. #StatisticsSaturday
34
34
1,216
156,681
codex completely broken for anyone else using tmux?
1
3
609
First CPU optimization from the naive implementation surprisingly really forces one to understand cache flow. Took me a non-trivial amount of time to realise it comes from the fact any rank R tensor still has to be stored contiguously. And the K-stride likely clobbers increasing cache levels as N grows. 8x speed up for N=2048.
65
Sasha retweeted
ai has made code incredibly cheap to produce, but not cheap to review if you spent 5 minutes generating a PR that takes someone else 2 hours to understand and validate, you've just moved the work downstream that's not productivity. you just made someone else do your job.
79
44
619
75,496
i'm getting nerdsniped by absolutely severything now that I have a superintelligence at my fingertips
2
345
holy shit people use GUIs? DataGrip is horrendous, how can anyone intuitively navigate this
470
Sasha retweeted
Someone asked me recently why I’ve become interested in aesthetics after having spent most of my life more interested in STEM-adjacent topics. I hadn’t really considered the question consciously before, but I’m certainly thinking about aesthetics more than I used to. I think it’s a confluence of things: • Many things today are ugly and far uglier than they used to be or need to be. Once you see this, it’s kinda hard to stop perceiving it. (Early twentieth century phone boxes versus modern phone boxes; old water fountains versus new water fountains; etc.) As someone with a naively meliorist assumption that most things should be getting better rather than worse, it’s all a bit perplexing: why did we stop doing things nicely? Is it a choice? Was there a malevolent spell cast upon us? This vein led me to think more about modernism and why much of art became more intentionally "challenging", grotesque, opposed to prettiness, rebarbative, dissonant, etc. Can or should anything be done about this? Is this just how things ought to be? • Relatedly, much of modernism involved a kind of explicit repudiation of cultural continuity and represented a schism with prior practices. This is maybe most evident in American architecture, where the International Style exhibition in 1932 initiated the displacement of a rich tapestry of prior styles. This cultural break seems important and interesting to me, and I suspect that the rejection had important consequences outside of the aesthetic domain. Samuel Hughes has been exploring this question in his recent writing at @WorksInProgMag; @RuxandraTeslo is also pulling on this thread. Elaine Scarry wrote about how beauty inspires creation. If so, the inverse may also be true: ugliness inhibits it. • It’s clearly the case that changes in the aesthetic domain can at least inspire progress in other places. Petrarch helped set some of the preconditions for the Renaissance which in turn fostered the scientific revolution and Enlightenment. Things like World Fairs (the 1851 Fair at the Crystal Palace recorded 6 million admissions when the population was 21 million) reflect the popular interdependency that used to exist between aesthetics and material development. • @tedgioia and others have written about stuck culture and how so many domains seem to have ceased to straightforwardly advance in the way that they did up until the nineties or thereabouts. This is obviously peculiar and interesting. What changed, and what does it mean? Is it about the internet and fragmentation? Is it about a loss of supply? Is it just about having reached the zenith of various mediums? • I’m generally interested in markets and the dynamics of creation. In aesthetics broadly, I find the reflexivity between supply- and demand-side factors to be very thought-provoking. There’s a natural desire to view satisfaction of individual preferences as the yardstick to measure market success, but things get interesting and even a bit unsettling when we start to think about how the supply might start to shape the demand. I often think about this in the context of food. Why is food so much worse in Germany than many of its neighbors? Germany certainly doesn’t have less material ability to produce good food; indeed, Germany is richer than the countries around it. There’s probably something about German food supply chains that is impoverished relative to France and Italy, but the Germans themselves don’t seem too upset about it. It just seems that the Germans are stuck in an objectively worse market equilibrium than their neighbors: the food is bad and they’ve gotten used to it. The obvious question then is where else these kinds of reflexive patterns apply, and where else we’re stuck in some objectively inferior equilibrium, even if preferences are in some superficial sense being sated. • While this is an extremely banal and obvious point, I hadn’t until recently thought much about or internalized how much one can study reasonably objective things ("the status of women in society", say) through artwork. (Thanks to @_alice_evans for opening my eyes here.) In this vein, I’m pretty excited about the possibilities over the coming years in computational art analysis. I want something that’s conceptually similar to Google Ngram timelines but for the visual arts. • We've always tried to do things well at Stripe. I've come to see that attempting to do them beautifully is often a helpful way to break out of standard practices and to do something with greater novelty and in a way that might have other benefits besides. (Also, excellent people want to do great work because it is intrinsically satisfying. Explicitly allowing aesthetic considerations to carry weight avoids having to justify every assessment with some kind of torturous empiricism.) • In his Nobel Lecture, Solzhenitsyn said that, among the Platonic virtues of goodness, truth, and beauty, that beauty is special, for it possesses a unique kind of irrefutability. He notes that arguments, writing, and philosophical systems can all be predicated on misapprehensions, but that “a true work of art carries its verification within itself.” He proceeds to observe that when goodness and truth are threatened, the “ever surprising shoots of beauty will still force their way through.” There is a lot of specious and motivated reasoning in the world today and plenty of questionable value systems. I don’t think that beauty directly reflects any definitive trait, but I’m intrigued by the idea that it can be a marker of deeper metaphysical coherence.
494
586
6,016
1,424,807
Sasha retweeted
I need my agents to turn every plan into a 3 blue 1 brown style animation so I can digest these without just staring at a giant ass wall of text
43
29
1,055
64,850