Official account for DeepSpeed, a library that enables unprecedented scale and speed for deep learning training + inference. 日本語 : @DeepSpeedAI_JP

Based in United States
Improved DeepNVMe: Affordable I/O Scaling for AI - Faster I/O with PCIe Gen5 - 20x faster model checkpointing - Low-budget SGLang inference via NVMe offloading - Pinned memory for CPU-only workloads - Zero-copy tensor type casting Blog: tinyurl.com/yanbrjy9
4
13
66
5,690
Introducing 🚀DeepCompile🚀: compiler-based distributed training optimizations. - Automatic parallelization & profile-guided optimizations - Enable ZeRO1, ZeRO3, Offloading, etc. via compiler passes - 1.2X-7X speedups over manual ZeRO1/ZeRO3/Offloading tinyurl.com/8cys28xk
1
51
304
42,627
AutoTP + ZeRO Training for HF Models - Enhance HF post-training with larger models, batches, & contexts - 4x faster LLAMA3 fine-tuning with TP=2 vs TP=1 - No code changes needed Blog: tinyurl.com/5n8nfs2w
19
75
10,371
🚀Introducing Ulysses-Offload🚀 - Unlock the power of long context LLM training and finetuning with our latest system optimizations - Train LLaMA3-8B on 2M tokens context using 4xA100-80GB - Achieve over 55% MFU Blog: shorturl.at/Spx6Y Tutorial: shorturl.at/bAWu5
1
28
96
5,806
Introducing Domino: a novel zero-cost communication tensor parallelism (TP) training engine for both single node and multi-node settings. - Near-complete communication hiding - Novel multi-node scalable TP solution Blog: github.com/microsoft/DeepSpe…
68
206
17,724
Announcing that DeepSpeed now runs natively on Windows. This exciting combination unlocks DeepSpeed optimizations to Windows users and empowers more people and organizations with AI innovations. - HF Inference & Finetuning - LoRA - CPU Offload Blog: shorturl.at/a7TF8
1
6
37
4,343
Introducing DeepNVMe, a suite of optimizations for fast and efficient I/O operations in DL applications. - POSIX-style APIs - Direct HBM/NVMe xfers via NVIDIA GDS - Cheap Inference scaling via NVMe-Offload Blog: shorturl.at/l7Oue @Azure @NVIDIADC #FMS24 #GPUDirect
16
54
16,580
Introducing Universal Checkpointing for boosting training efficiency. - Change parallelism (PP, SP, TP, ZeRO-DP) or GPU count mid-stream - Improve resilience by scaling down to healthy nodes💪 - Increase throughput by scaling up to elastic nodes🚀 Blog: rb.gy/aup3pn
5
23
4,291
#DeepSpeed joins forces with @Sydney_Uni to unveil an exciting tech #FP6. Just supply your FP16 models, and we deliver: 🚀 1.5x performance boost for #LLMs serving on #GPUs 🚀 Innovative (4+2)-bit system design 🚀 Quality-preserving quantization link: github.com/microsoft/DeepSpe…
26
166
18,953
Introducing Mixtral, Phi2, Falcon, and Qwen support in #DeepSpeed-FastGen! - Up to 2.5x faster LLM inference - Optimized SplitFuse and token sampling - Exciting new features like RESTful API and more! For more details: github.com/microsoft/DeepSpe… #DeepSpeeed #AI
9
87
409
49,560
🚀 Announcing DeepSpeed ZeRO-Offload++ -6x Higher Training Throughput via Collaborative CPU/GPU Twin-Flow 🔥 -Systematic optimizations at no data precision loss -Performance gain maintains for both single and multi-node cases github.com/microsoft/DeepSpe…
7
60
291
50,982
Introducing DeepSpeed-FastGen 🚀 Serve LLMs and generative AI models with - 2.3x higher throughput - 2x lower average latency - 4x lower tail latency w. Dynamic SplitFuse batching Auto TP, load balancing w. perfect linear scaling, plus easy-to-use API github.com/microsoft/DeepSpe…
6
113
542
112,905
🚀Introducing #DeepSpeed-VisualChat! 🖼📜 - Multi-image, multi-round #dialogues - Novel #MultiModal causal attention - Enriched training data via improved blending techniques - Unmatched #scalability (>70B params) Blog: github.com/microsoft/DeepSpe… Paper: arxiv.org/abs/2309.14327
1
35
133
18,527
Highlight 2: Scientists can now train their large science models like Argonne's GenSLM COVID models with very long sequences - 2X higher training throughput 🚀 - 13X longer sequence lengths achieved compared to SOTA training frameworks like Megatron-LM deepspeed.ai/deepspeed4scien…
Announcing DeepSpeed4Science 🚀 We are building unique capabilities through AI system technologies to help domain experts solve society's most pressing science challenges, from drug design to renewable energy. MSR Blog: bit.ly/48iTGem Website: deepspeed4science.ai/
4
8
1,373
Highlight 1: Eliminating memory explosion problems for scaling Evoformer-centric structural biology models🧬 Today we’re releasing a set of highly memory-efficient Evoformer attention kernels that reduces peak memory for training and inference by 13x!🚀 deepspeed.ai/deepspeed4scien…
Announcing DeepSpeed4Science 🚀 We are building unique capabilities through AI system technologies to help domain experts solve society's most pressing science challenges, from drug design to renewable energy. MSR Blog: bit.ly/48iTGem Website: deepspeed4science.ai/
1
2
8
1,119
Announcing DeepSpeed4Science 🚀 We are building unique capabilities through AI system technologies to help domain experts solve society's most pressing science challenges, from drug design to renewable energy. MSR Blog: bit.ly/48iTGem Website: deepspeed4science.ai/
8
36
5,633
🚀Exciting new updates on #DeepSpeed ZeRO-Inference with 20X faster generation! - 4x lesser memory usage through 4-bit weight quantization with no code change needed. - 4x larger batch sizes through KV cache offloading. Available in DeepSpeed v0.10.3: aka.ms/z3-inference
2
28
166
18,166
🚀 Exciting Updates for #DeepSpeedChat! 🤖 - Llama-2 Support: Enjoy 7.1x faster generation with DeepSpeed Hybrid Engine! - Improved efficiency and accessibility through MixZ++ and ZeRO-Offload. - Improved stability and software enhancements. Blog: github.com/microsoft/DeepSpe…
2
18
94
5,844
Want to train 1 million token context lengths (all 7 of the Harry Potter books!📚) on a GPT-like model w. 64 GPUs? Announcing DeepSpeed-Ulysses🚀 This release enables highly efficient and scalable LLM training with extremely long sequence lengths🤯 github.com/microsoft/DeepSpe…
1
39
139
15,745
Want to train 10B+ ChatGPT-style models on a single GPU and 100B+ on multi-GPUs systems? Introducing DeepSpeed-Chat, an easy (single script), fast, and low-cost solution for training high-quality ChatGPT-style models with RLHF, 15x faster than SoTA. Blog: github.com/microsoft/DeepSpe…
13
130
437
167,377
Scaling Large-Scale Generative Mixture-of-Expert Multimodal Model With VL-MoE Do you want to scale up your vision and language models? Take a look at our blog for details! deepspeed.ai/2023/03/30/mult…
DeepSpeed + @berkeley_ai explore the effectiveness of MoE in scaling vision-language models, demonstrating its potential to achieve state-of-the-art performance on a range of benchmarks over dense models w. equivalent compute costs. arxiv.org/abs/2303.07226 More coming soon!
7
15
1,711
DeepSpeed + @berkeley_ai explore the effectiveness of MoE in scaling vision-language models, demonstrating its potential to achieve state-of-the-art performance on a range of benchmarks over dense models w. equivalent compute costs. arxiv.org/abs/2303.07226 More coming soon!
11
28
4,409