Data parallelism
Two-GPU data parallelism Forward sends different batch shards through identical replicas. Backward averages local gradients using AllReduce. global batch B₀ B₁ GPU 0 replica layer 1 layer 2 layer 3 GPU 1 replica layer 1 layer 2 layer 3 Loss L₀ Loss L₁ no cross-GPU communication local gradients g₀ g₁ ALL-REDUCE