OK last post for the night: I tried all the fancy stuff they recommended in their GEMM doc: Z-curve, static extents, accumulation group synchronization. None of it seemed to make any performance improvement - I seem to be stuck at 40 TFLOPs in bf16 across a variety of shapes.
I got my M5 MacBook over the weekend and had some time to mess around with Metal 4 and the Neural Accelerators!
Wanted to document some of my first impressions below: