Of course, that’s your contention. You’re a first-year machine-learning engineer. You just got finished readin' Attention Is All You Need, probably watched a Karpathy video too. So now you’re convinced everything is just transformers and scaling laws.
That's gonna last until next month when you discover convolution, and then you're gonna be talkin' about inductive biases and locality and how CNNs were actually incredibly compute-efficient for vision.
Then you're gonna read the FlashAttention paper, and suddenly everything's about IO complexity and SRAM and how FLOPs don't matter because the whole goddamn thing is memory-bandwidth bound.
That'll last until somebody shows you an MoE model, and then you're gonna be in here regurgitating DeepSeek, talkin' about expert parallelism and active parameters and how dense models are economically obsolete.
Then six months from now you'll write one shitty Triton kernel, look at an Nsight trace for the first time, and start telling everybody Python isn't actually the bottleneck because the GPU is asynchronous.
And by next year you're gonna be standing right here explaining to me that your 400-billion-parameter model is fast because you turned on CUDA graphs.
Ben Affleck reveals he writes Python, understands convolutional neural networks, worked extensively with GPUs, and used his celebrity status to get private looks at Google and OpenAI’s video models
“I’ve always been kind of into computers since I was young. Then, when film started to move from analog film to digital, I became more interested in that aspect of it. The visual-effects workflow for many years has included machine learning, so I can write pretty shitty Python scripts and stuff like that.
“With convolutional neural networks, which were the precursors to what the transformer can do, which is much more computation simultaneously, you would do things like look at what’s called a tensor. That’s the numerical translation of a visual image in numbers, like the batch number, the frame number and the red, green and blue values of each pixel in each frame. It’s just that simple. That numeric is called a tensor.
“You’d use a convolutional neural network to identify patterns that reveal what’s called edge detection or feature extraction, which is identifying patterns well enough to know, this is where the window ledge is, so we can more easily take the green-screen image out and replace it with something.
“That was familiar to me early on because, prior to Artists Equity, I had a small visual-effects company. I’ve worked with GPUs a lot too. The visual-effects guys said, ‘Hey, you should see. There are a couple: Google and this other company, OpenAI, are doing really interesting stuff with transformers in video.’
“I’ve learned that I can actually just call up and go, ‘Hey, it’s Ben Affleck. Can I come see what you’re doing?’ Sometimes people say yes, to my astonishment.”