Important result - distributed training of Chinchilla-style LLMs out to 10B parameters gets you a LOWER loss than standard data parallel training if training across two replicas (M=2). For AI policy, the key question is if M can go above something like ~10.
10
18
204
27,132
Replying to @jackclarkSF

Mar 16, 2025 · 12:16 AM UTC

163
Sort replies: Relevant Recent Liked