Expert parallelism top-2 / 2 EP ranks / 4 experts per rank
Vertical top-2 expert routing across two ranks Blue token x0 and green token x1 enter replicated routers from the top. Each token is sent to one local and one remote expert while retaining its color. Expert outputs return to the token's source rank and are combined at the bottom. source rank 0 source rank 1 local token local token x0 x1 input grad input grad dx0 dx1 replicated router replicated router scores global experts E0-E7 scores global experts E0-E7 top-2: E1, E6 top-2: E2, E5 DISPATCH ALL-TO-ALL solid = local / dashed = crosses rank INVERSE DISPATCH ALL-TO-ALL return input-gradient contributions EP rank 0 / owns E0-E3 EP rank 1 / owns E4-E7 E0 E1 E2 E3 x0 x1 E4 E5 E6 E7 x1 x0 expert gradients stay on rank 0 expert gradients stay on rank 1 RETURN ALL-TO-ALL expert outputs return to token owners INVERSE RETURN ALL-TO-ALL route output gradients to expert owners weighted combine weighted combine y0 y1 dy0 dy1 arrow color follows the token through local and remote experts ALL-REDUCE router + shared grads backward reverses both exchanges; sharded expert gradients stay with their owners