AI explained
What is attention, and what is it for?
Follow an ambiguous word through queries, keys and values to see how attention combines information from context, and what its weights can tell us.
Article 4 of 5 · Reading order
Imagine finding the word bank on a page. It could be somewhere to keep money or the edge of a river. Now read a little further: “The fisherman repaired his net beside the bank.” Replace that sentence with “The cashier counted the money inside the bank,” and a different scene appears.

The same word, different surrounding clues: a conceptual illustration of context, not a measured attention map.
For us, the surrounding words offer clues that the word alone could not provide. For a model, those clues have to participate in a calculation. We met bank and a fisherman in the article on tokens and embeddings, where we saw how an identifier retrieves a vector from a table. If both occurrences of bank correspond to the same token, they retrieve the same row. What we need now is a way to incorporate information from this particular sentence.
Attention combines information from different positions, giving each contribution a weight that depends on the query. We will follow that operation from the position of bank: where the numbers being compared come from, how they become weights and what happens when we use them. By the end, we will be able to describe something more concrete than “the model focuses on what matters.”
A word needs information from its neighbors
Think of a representation as a list of numbers that a layer receives and transforms. Initially, it may come from an embedding and positional information; further along, it may contain information other layers have already combined. We want to see how something contributed by fisherman and net reaches the position of bank.
One possibility would be to average all the vectors equally. But that rule would deliver the same mixture to every position. The representations of bank, repaired and net might need different combinations. Attention lets us calculate a mixture for each query.
That need had an influential application in translation. Bahdanau, Cho and Bengio proposed allowing a system to combine representations from different parts of the source sentence as it generated a word of the translation. Their work circulated in 2014 and was presented at ICLR 2015. It provided a different context at each step, instead of relying solely on a fixed summary of the entire input (Bahdanau et al., 2015 (opens in a new tab)).
In our example, the positions exchanging information belong to the same sentence. This is called self-attention. We will follow the scaled dot-product version used by the Transformer. Bahdanau and colleagues used additive attention, which scores compatibility with a small feedforward network. Both turn scores into a weighted mixture; they calculate the scores differently. Attention can also connect two different sequences, as we will see when we return to translation.
Three roles for the numbers: query, key and value
Imagine looking for a note in a filing system. You have a search criterion, something to compare it with and content to retrieve. That distinction helps introduce three roles in attention: query, key and value, usually abbreviated Q, K and V. The analogy has a limit: here everything is a vector, and the result can blend several contributions instead of retrieving a single record.
In the self-attention operation we are following, every position produces all three vectors from its representation. The query takes part in comparisons initiated from that position. The key takes part when a query is compared with it. The value contains the numbers that will enter the mixture. Keys help calculate how much each position contributes; values supply the numerical content.
To obtain them, we multiply the input representation by three matrices of learned parameters. If is the row vector this layer receives at position :
Each matrix combines the input coordinates using its own coefficients. Query, key and value can therefore differ even though they start from the same representation. Within an attention head, the same matrices are applied to every position. This construction, and the separation between comparison and content, appear in the Transformer's attention mechanism (Vaswani et al., 2017 (opens in a new tab), section 3.2 (opens in a new tab)).
For our sentence, we will follow the query at bank. “River edge or financial institution?” describes our reading problem; the model's query is a vector. It will be compared with the keys at the available positions, including its own.
From a comparison to a proportion
We need to turn each comparison into a number. We return to the dot product used in the previous article: multiply corresponding coordinates of two vectors and add the results. For a query and key with two dimensions, that would be . In the version we are following, we also divide by the square root of the key's number of dimensions:
Here is the querying position, the position being queried and the number of coordinates in each query and key. The division moderates the scale of the scores; note 1 explains why. Unlike the cosine we used to explore embeddings, this calculation does not divide by the lengths of the two vectors: their magnitudes matter too.
Let us work through an example we can check by hand. All the numbers are invented. We will assume one token per word and reduce the calculation to three positions, fisherman, net and bank, so we can see the whole operation with three available contributions.
Choose the query for bank and these keys, with :
| Token | ||
|---|---|---|
fisherman |
2 | |
net |
1 | |
bank |
0 |
For example, the dot product with the key at fisherman is . Dividing by gives the score of 2 in the table.
The scores are 2, 1 and 0. Now we want positive proportions that add up to one so we can use them to weight a mixture. Softmax raises the constant — approximately 2.718 — to each score and divides each result by their total:
The displayed values are rounded; before rounding, they sum to one. A score of zero does not disappear either: . In this exercise, fisherman receives the largest share, but all three positions contribute. Differences between scores determine how the weight is distributed across the values (Vaswani et al., 2017 (opens in a new tab), section 3.2.1 (opens in a new tab)).
That 0.6652 does not mean “there is a 66.52% probability that bank means a river edge.” It describes how strongly one of the values will be weighted in this calculation. Its size depends on this query, these keys and the positions competing with it.
Values are what travel
We now know how much each position will contribute. We still need the numbers the operation will combine. That is what values are for. Assign to fisherman, to net and to bank, with no linguistic labels attached to their coordinates.
We multiply each vector by its weight and add coordinate by coordinate:
The final result uses the unrounded weights. The first coordinate draws mainly on the contribution we assigned to fisherman; the second receives a larger contribution from net. Equal weights would have produced . The difference lets us see what weighting the contributions changed.

One head’s complete calculation for one position, reduced to three invented contributions. Bars represent mixture weights; the result uses unrounded weights.
This is the output of one attention head for one position. We have built a representation using information from other positions, but we have not yet produced the words “river edge” or a classification. Later layers can use the result for the model's task. Note 2 shows how to express the same procedure for all positions using matrices.
Now replace the fisherman's scene with the cashier's. The surrounding representations change, and so do their keys and values. If we keep bank as the same token at the same position in both sentences, its initial representation and first-layer query remain the same in this setup. That query can still receive a different mixture from its changed neighbors. In later layers, the query at bank can change too, because its representation has already incorporated context. Thus the same row in the embedding table can lead to different contextual representations.
What the model learns and what it calculates as it reads
It may seem that we chose just the right numbers to get the mixture we wanted. That is exactly what we did to explain the calculation. During model training, however, the matrices , and are adjusted through the task's error signal, together with the other parameters. The article on neural networks develops that cycle of prediction and adjustment.
It helps to distinguish two uses of the word weight. The coefficients in those matrices are learned parameters, which ordinary inference keeps fixed. Attention weights, such as 0.6652, are results calculated for a particular input. They can change when the text changes even though the model's parameters remain the same.
That distinction takes us back to the opening scene: the model can process a sentence about a river and then one about money without being retrained between them. It applies the transformations it learned to different inputs. Training aims to make those transformations useful; tests on new examples let us evaluate how far that usefulness extends.
Several mixtures for the same word
Our mixture gave fisherman a weight of about 0.6652 and net about 0.2447. The full-precision weights sum to one: giving more weight to one position leaves less for the others. If we wanted a second mixture that emphasized net, changing this distribution would also change the result we have just calculated.
A second head has its own , and . Suppose its queries and keys produce scores for the same three positions: we have deliberately swapped the scores for fisherman and net. Softmax swaps their weights too, so net now receives about 0.6652 in this head, while the first head keeps its original mixture. Each head distributes its own total of one. The two mixtures are available together, and each can carry information through its own value projection.
Multi-head attention concatenates those results and applies an output projection to combine them (Vaswani et al., 2017 (opens in a new tab), section 3.2.2 (opens in a new tab)). The usefulness is in retaining several differently weighted combinations for that next step. Which relationships each head carries emerges from learning and can be investigated in the trained model.

Two independently normalized mixtures, using the invented scores from our example. Each head's weights sum to one before rounding; their outputs are concatenated and projected.
Which positions can contribute?
The previous article introduced the causal restriction in language models: a position can use itself and earlier positions. We can now locate that restriction inside the attention calculation. A mask excludes later positions before softmax, giving them zero weight. Our fisherman and net precede bank, so both remain available.
Return to “The bank flooded after the storm.” Under that causal mask, the query at bank cannot use the key or value at the later storm position. A query at storm can combine information from both. The previous article also introduced bidirectional models, which allow context from both sides. In bidirectional self-attention, bank can access the later clue. Even then, the sentence could describe a flooded riverbank or bank branch: if the clues are insufficient, the model still has a problem to resolve.

Available positions for the same query. The marks show access, not attention weights. We assume one token per word and a separate token for the period.
In cross-attention, queries come from one sequence while keys and values come from another. In translation, an output representation can query input representations (Vaswani et al., 2017 (opens in a new tab), section 3.2.3 (opens in a new tab)). This returns us to Bahdanau and colleagues: their translation mechanism also connected the developing output with representations of the source sentence, using additive scoring. The source of the information changes; the idea of calculating a mixture for a query remains.
An attention map invites us to look, and to test
Color each weight by its intensity and we get a map that is easy to inspect. It would be tempting to point at the darkest cell and say, “Here is the reason for the answer.” Our calculation already reveals a difficulty: the weight is only part of the contribution; the vector it multiplies and what the rest of the model does with the result also matter.
There is direct evidence for looking beyond weights in Transformers. Kobayashi and colleagues analyzed BERT and a Transformer translation system using both attention weights and the norms of the transformed input vectors, including the value and output projections. Here, a norm measures a vector's length: a large weight multiplying a short vector can contribute less than the weight alone suggests. In the BERT models they studied, special tokens such as the separator [SEP] contributed much less to the attention output than their large weights suggested, because their transformed vectors had small norms (Kobayashi et al., 2020 (opens in a new tab)).
A related debate studied a different architecture. In experiments centered on bidirectional recurrent encoders (BiLSTMs) with a single attention layer, Jain and Wallace constructed, through optimization, very different attention distributions yielding essentially unchanged predictions in classification, question answering and natural language inference (Jain & Wallace, 2019 (opens in a new tab)). Wiegreffe and Pinter revisited LSTM classifiers and proposed further tests, including whether a trained model could produce the alternative distributions (Wiegreffe & Pinter, 2019 (opens in a new tab)). Those results do not settle the interpretability of multi-head self-attention; they show why the architecture, the intervention and the meaning of “explanation” belong in the question.
For a trained model processing our sentence, a map could suggest a relationship to investigate. To test whether information from fisherman helps resolve bank, we could change clues in the text or intervene in a head's contribution, then measure how the output changes. The map would help formulate a hypothesis; the experiment would test it.
Back to the riverbank with an operation we can follow
We began with a word that allowed two scenes. Now we can follow one route by which context participates in its representation: a query is compared with keys, scores become weights and those weights combine values. The result brings information from other positions to one position, in proportions that depend on the text and the learned transformations.
The fisherman's sentence helps us picture what this is for; the small calculation lets us see how it happens. We can now ask which clues are available, how their contributions are mixed and what changes when we alter them.
We still need to understand how this operation fits into a complete network, how information is preserved between layers and how the sentence's order is represented. That will be our route through the Transformer. For now, when we read that a model “pays attention,” we can picture something concrete: numbers that compare, weight and carry information.
Notes
-
Why divide by a square root? Under the illustrative assumption that query and key components are independent, with mean zero and variance one, their dot product has variance . Dividing by keeps that variance at one. This helps prevent scores from growing with dimensionality and pushing softmax into regions with very small gradients. It is the motivation for scaling, not a guarantee that learned vectors satisfy those assumptions (Vaswani et al., 2017 (opens in a new tab), section 3.2.1 (opens in a new tab)). Return to the scores.
-
The same calculation in matrices. Let contain the representations of positions, with coordinates each. With and , we obtain , and . Thus and . The score matrix has shape . Applying softmax to each row and multiplying by gives : one mixture per query. A causal mask adds to disallowed scores before softmax so they receive zero weight. The formula describes one head; the values' dimensionality, , may differ from (Vaswani et al., 2017 (opens in a new tab), section 3.2 (opens in a new tab)). Return to the values.
References
Bahdanau, D., Cho, K., & Bengio, Y. (2015, May 7–9). Neural machine translation by jointly learning to align and translate [Paper presentation]. 3rd International Conference on Learning Representations, San Diego, CA, United States. Conference-era manuscript, arXiv v6 (April 24, 2015) (opens in a new tab)
Jain, S., & Wallace, B. C. (2019). Attention is not explanation. In J. Burstein, C. Doran, & T. Solorio (Eds.), Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) (pp. 3543–3556). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1357 (opens in a new tab)
Kobayashi, G., Kuribayashi, T., Yokoi, S., & Inui, K. (2020). Attention is not only a weight: Analyzing Transformers with vector norms. In B. Webber, T. Cohn, Y. He, & Y. Liu (Eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 7057–7075). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.emnlp-main.574 (opens in a new tab)
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In I. Guyon, U. von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, & R. Garnett (Eds.), Advances in Neural Information Processing Systems (Vol. 30, pp. 5998–6008). Curran Associates. Conference proceedings (opens in a new tab)
Wiegreffe, S., & Pinter, Y. (2019). Attention is not not explanation. In K. Inui, J. Jiang, V. Ng, & X. Wan (Eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 11–20). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1002 (opens in a new tab)