The exact operation at the heart of every modern language model. Each token forms a query, key and value; token i attends to token j with weight softmax(qᵢ·kⱼ/√d_k). Multiple heads attend in different learned subspaces, then their outputs are concatenated and projected. Everything below is computed live from the tokens you type.