dgpp, a C++/CUDA engine for DGX Spark, developed a system where each rank keeps its model slice resident on the GPU and boots from a per rank image cache in 15–30s, depending on the model.
I just looked at adopting it, but our EXL3/DFlash2 and deepseek recipes aren't supported so far.
Watching for compatible support...