Expert parallelism
top-2 / 2 EP ranks / 4 experts per rank
Vertical top-2 expert routing across two ranks
Blue token x0 and green token x1 enter replicated routers from the top. Each token is sent to one local and one remote expert while retaining its color. Expert outputs return to the token's source rank and are combined at the bottom.
source rank 0
source rank 1
local token
local token
x0
x1
input grad
input grad
dx0
dx1
replicated router
replicated router
scores global experts E0-E7
scores global experts E0-E7
top-2: E1, E6
top-2: E2, E5
DISPATCH ALL-TO-ALL
solid = local / dashed = crosses rank
INVERSE DISPATCH ALL-TO-ALL
return input-gradient contributions
EP rank 0 / owns E0-E3
EP rank 1 / owns E4-E7
E0
E1
E2
E3
x0
x1
E4
E5
E6
E7
x1
x0
expert gradients stay on rank 0
expert gradients stay on rank 1
RETURN ALL-TO-ALL
expert outputs return to token owners
INVERSE RETURN ALL-TO-ALL
route output gradients to expert owners
weighted combine
weighted combine
y0
y1
dy0
dy1
arrow color follows the token through local and remote experts
ALL-REDUCE
router + shared grads
backward reverses both exchanges; sharded expert gradients stay with their owners
forward
backward