Multi-vendor is a means, not an achievement. It matters because a run can then be assembled from whatever hardware is around rather than whatever is identical, and uneven hardware is what flodl's distributed layer was built for. Next: real AMD silicon. flodl.dev/blog/making-room
1
The Mac is the argument for that. flodl had never actually run on one: it compiled there perfectly well and failed at launch, for reasons only running it could surface. Fixed, 100 epochs in about 1.5s on a MacBook, and the test is a hard failure now rather than a warning.
8
Where it runs: seven platforms on every push. Two Ubuntu, Ubuntu on arm, macOS, Windows, two Rocky. Each installs the GPU toolkit the way the docs tell a reader to install it, natively, no container smoothing it over. The install instructions are tested, not just written.
2
There is no AMD card in this workshop, so the numbers that would tell you whether any of it is good do not exist yet. I would rather say that here than let someone find out after an afternoon of setup. When they arrive they arrive as numbers, not as a claim.
1
flodl 0.8.0 supports a second GPU vendor. AMD cards are detected, the code builds and links against ROCm, the whole test suite runs against it. No AMD card has ever run a training step. Both are true. The second one is the more useful. 🧵 flodl.dev/blog/making-room
6
The published sweep, re-run seeded: ResNet-20 on CIFAR-10 across three mismatched GPUs on two hosts. DiLoCo over async CPU averaging hits 92.29% against the 91.25% reference, and finishes 28% sooner than the fast GPU alone.
26
Same record plane, four ways: the live portal, an HTTP path API (/paths, /node, /history, /stream), a JSONL tree where a record's path IS its file path, and the whole thing saved as one self-contained HTML file. No server needed to read it later.
24
flodl 0.7.0: the training dashboard is now a recursive portal. One view, repeated at every level of a cluster, root to host to rank. The page subscribes to one level, so a 300-rank run costs the same to watch as a 3-rank one. 🧵
23
The surprise: solo training reached that exact 92.36% at epoch 159, then overfit back down to 92.10%. In the distributed consensus, replica-private memorization gets averaged away every round. DiLoCo holds the peak. Distributed training as built-in early stopping.
29
The flagship: ResNet-20 on CIFAR-10, 200 epochs. RTX 5060 Ti + two GTX 1060 (one behind a PCIe x1 riser, ~12x slower delivered). The cluster: 92.36% in 524s. The fast card alone: 92.10% in 699s. Published reference: 91.25%. Same seed, same code, one config line per mode.
1
45
Three months holding my breath after three months of weekly releases. Not by choice. A dead CPU forced a rig rebuild, then forced a VM, and GPUs in a VM are GPUs on another machine. The thread-based path stopped existing overnight. All or nothing: 366 commits later, here it is.
1
26
My CPU died in April. flodl 0.6.0 is what it cost me: the Rust deep learning framework went from threads on one box to full multi-host DDP. Three mismatched GPUs across two hosts now beat the fastest card alone, on accuracy AND wall time. 🧵 ResNet-20 on CIFAR-10
37
A declarative graph DSL for neural networks. Tag streams, build residuals, compose trained subgraphs as frozen Modules in larger architectures. Every level addressable by name. "The structure IS the text." Full post: dev.to/fab2s/i-wanted-to-des…
30
Phase one was loading. Phase two was fine-tuning, now plumbed end-to-end. Phase three is the demo on the hardware people actually own. Next: ModernBERT, LLaMA, LoRA, ViT, plus a flagship El Che fine-tune benchmark on heterogeneous consumer GPUs.
43
Three new families: ALBERT (factorised embeddings, cross-layer sharing) XLM-RoBERTa (multilingual SentencePiece, ~250k vocab) DeBERTa-v2/v3 (disentangled attention, mask-gated embeddings) MLM heads across all six families. fill_mask one-call for any checkpoint.
72
Universal Trainer for transparent fine-tuning. You write one closure (forward + loss); the framework owns the loop, backward, optimizer step, gradient sync. Same code on CPU, single GPU, or heterogeneous multi-GPU. El Che cadence auto-tunes the slow card.
29
Round-trip export is the flagship. fdl flodl-hf export --hub <repo> --out staged/ fdl flodl-hf verify-export staged/ Re-emits any flodl-hf checkpoint as an HF-canonical dir that loads back into HF Python's AutoModelFor* with bit-exact agreement on every head output.
23
🧵 flodl 0.5.3: HuggingFace, both ways. 30 head cells, bit-exact round-trip with HF Python (max abs diff = 0). 6 BERT-family architectures × 5 head shapes. In Rust. flodl.dev/blog/huggingface-b… #rustlang #huggingface #deeplearning
1
38