P₃
O₂
G₂
P₂
O₁
~25% Memory
G₁
P₁
GPU 3
GPU 2
ZeRO Stage 3
GPU 1
~50% Memory
ZeRO Stage 2
All-Reduce Communication
💾 Memory: 1/4 of Full
O₄
Data Flow
G₄
P₄
Optimizer States (O)
O₃
G₃
Gradients (G)
Parameters (P)
Legend
显存占用对比 (Memory Usage Comparison)
Memory Comparison
💡 Key Insight: ZeRO通过分区策略将显存占用从100%降低至25%,支持更大模型训练
DeepSpeed ZeRO 分布式训练架构
Input Layer
~62.5% Memory
Data Batch
📦 Training Data Samples
ZeRO Stage 1
ZeRO Partitioning Strategy
State Partitioning - 每个GPU仅保存部分状态
GPU 0
100% Memory
Standard DP