Multi-Head Attention
Head 1
Q1
K1
Head 2
Q2
K2
Updated "woods" embedding
Fig. 6. Each head has its own Q and K. Head 1 writes the top half of the embedding; Head 2 writes the bottom half — then they concatenate.