Can a model run deeper at test time than it was ever trained to? And if depth becomes a loop instead of a stack, do we need better optimizers to keep training stable?
@coreauto is collaborating with
@tilderesearch on an optimizer x architecture competition: "One Layer Deeper"
We picked a cursed problem y = x^(2^T) mod N where each squaring depends on the last. A vanilla transformer has a fixed depth so it falls flat on its face the moment T exceeds the serial compute it can do in one forward pass. A fun reminder that architectures have inductive biases.
This isn't quite a nano-gpt speedrun experience and it's not a typical
@GPU_MODE kernel competition either, it's closer in spirit to the LLM efficiency competition I worked on many years ago with
@weiwei_msr and
@Jisaacso in that it's open ended in a specific way.
The UX is kernelbot like, your submission is a single python file that defines your model, optimizer and loss. You don't need a GPU to participate at all since we're reusing a queue based job system powered by
@modal and
@northflank
There's tons of unexplored ideas and I'm really not sure what will win out but I'd be particularly excited to see people playing with new optimizers, adaptive depth and depth extrapolation.
Since this is a new format, we'll work with the community to refine the rules for another week before we start for real. Submissions are open now!!