AutoTP + ZeRO Training for HF Models
- Enhance HF post-training with larger models, batches, & contexts
- 4x faster LLAMA3 fine-tuning with TP=2 vs TP=1
- No code changes needed
Blog: tinyurl.com/5n8nfs2w
🚀Introducing Ulysses-Offload🚀
- Unlock the power of long context LLM training and finetuning with our latest system optimizations
- Train LLaMA3-8B on 2M tokens context using 4xA100-80GB
- Achieve over 55% MFU
Blog: shorturl.at/Spx6Y
Tutorial: shorturl.at/bAWu5
Introducing Domino: a novel zero-cost communication tensor parallelism (TP) training engine for both single node and multi-node settings.
- Near-complete communication hiding
- Novel multi-node scalable TP solution
Blog: github.com/microsoft/DeepSpe…
Announcing that DeepSpeed now runs natively on Windows. This exciting combination unlocks DeepSpeed optimizations to Windows users and empowers more people and organizations with AI innovations.
- HF Inference & Finetuning
- LoRA
- CPU Offload
Blog: shorturl.at/a7TF8
Introducing DeepNVMe, a suite of optimizations for fast and efficient I/O operations in DL applications.
- POSIX-style APIs
- Direct HBM/NVMe xfers via NVIDIA GDS
- Cheap Inference scaling via NVMe-Offload
Blog: shorturl.at/l7Oue@Azure@NVIDIADC#FMS24#GPUDirect
Introducing Universal Checkpointing for boosting training efficiency.
- Change parallelism (PP, SP, TP, ZeRO-DP) or GPU count mid-stream
- Improve resilience by scaling down to healthy nodes💪
- Increase throughput by scaling up to elastic nodes🚀
Blog: rb.gy/aup3pn
#DeepSpeed joins forces with @Sydney_Uni to unveil an exciting tech #FP6. Just supply your FP16 models, and we deliver:
🚀 1.5x performance boost for #LLMs serving on #GPUs
🚀 Innovative (4+2)-bit system design
🚀 Quality-preserving quantization
link: github.com/microsoft/DeepSpe…
Introducing Mixtral, Phi2, Falcon, and Qwen support in #DeepSpeed-FastGen!
- Up to 2.5x faster LLM inference
- Optimized SplitFuse and token sampling
- Exciting new features like RESTful API and more!
For more details: github.com/microsoft/DeepSpe…#DeepSpeeed#AI
🚀 Announcing DeepSpeed ZeRO-Offload++
-6x Higher Training Throughput via Collaborative CPU/GPU Twin-Flow 🔥
-Systematic optimizations at no data precision loss
-Performance gain maintains for both single and multi-node cases
github.com/microsoft/DeepSpe…
Introducing DeepSpeed-FastGen 🚀
Serve LLMs and generative AI models with
- 2.3x higher throughput
- 2x lower average latency
- 4x lower tail latency
w. Dynamic SplitFuse batching
Auto TP, load balancing w. perfect linear scaling, plus easy-to-use API
github.com/microsoft/DeepSpe…
Highlight 2:
Scientists can now train their large science models like Argonne's GenSLM COVID models with very long sequences
- 2X higher training throughput 🚀
- 13X longer sequence lengths achieved compared to SOTA training frameworks like Megatron-LM
deepspeed.ai/deepspeed4scien…
Announcing DeepSpeed4Science 🚀
We are building unique capabilities through AI system technologies to help domain experts solve society's most pressing science challenges, from drug design to renewable energy.
MSR Blog: bit.ly/48iTGem
Website: deepspeed4science.ai/
Highlight 1:
Eliminating memory explosion problems for scaling Evoformer-centric structural biology models🧬
Today we’re releasing a set of highly memory-efficient Evoformer attention kernels that reduces peak memory for training and inference by 13x!🚀
deepspeed.ai/deepspeed4scien…
Announcing DeepSpeed4Science 🚀
We are building unique capabilities through AI system technologies to help domain experts solve society's most pressing science challenges, from drug design to renewable energy.
MSR Blog: bit.ly/48iTGem
Website: deepspeed4science.ai/
Announcing DeepSpeed4Science 🚀
We are building unique capabilities through AI system technologies to help domain experts solve society's most pressing science challenges, from drug design to renewable energy.
MSR Blog: bit.ly/48iTGem
Website: deepspeed4science.ai/
🚀Exciting new updates on #DeepSpeed ZeRO-Inference with 20X faster generation!
- 4x lesser memory usage through 4-bit weight quantization with no code change needed.
- 4x larger batch sizes through KV cache offloading.
Available in DeepSpeed v0.10.3: aka.ms/z3-inference
Want to train 1 million token context lengths (all 7 of the Harry Potter books!📚) on a GPT-like model w. 64 GPUs?
Announcing DeepSpeed-Ulysses🚀
This release enables highly efficient and scalable LLM training with extremely long sequence lengths🤯
github.com/microsoft/DeepSpe…
Want to train 10B+ ChatGPT-style models on a single GPU and 100B+ on multi-GPUs systems? Introducing DeepSpeed-Chat, an easy (single script), fast, and low-cost solution for training high-quality ChatGPT-style models with RLHF, 15x faster than SoTA.
Blog: github.com/microsoft/DeepSpe…
ALT DeepSpeed-Chat: Easy, Fast & Low-Cost Training of ChatGPT-like Models
Scaling Large-Scale Generative Mixture-of-Expert Multimodal Model With VL-MoE
Do you want to scale up your vision and language models? Take a look at our blog for details!
deepspeed.ai/2023/03/30/mult…
ALT New encoding process in our VL-MoE for various modality inputs, for which gray and colored blocks indicate non-activated and activated modules, respectively.
DeepSpeed + @berkeley_ai explore the effectiveness of MoE in scaling vision-language models, demonstrating its potential to achieve state-of-the-art performance on a range of benchmarks over dense models w. equivalent compute costs.
arxiv.org/abs/2303.07226
More coming soon!
DeepSpeed + @berkeley_ai explore the effectiveness of MoE in scaling vision-language models, demonstrating its potential to achieve state-of-the-art performance on a range of benchmarks over dense models w. equivalent compute costs.
arxiv.org/abs/2303.07226
More coming soon!