AI explained
What is a Transformer?
Follow the fisherman from an English sentence to its translation and discover how a Transformer connects information to build an output, without needing formulas.
Article 5 of 5 · Reading order
Imagine translating this sentence: “The fisherman repaired his net beside the bank.” If you have followed the series, you already know the fisherman repairing his net by the river. The English word bank helped us see why context matters. Now we will follow it all the way to a translation.

One scene, several transformations. The fisherman will stay with us throughout the journey.
The difficulty goes beyond picking the Spanish word orilla from a dictionary. We need to preserve who repairs what, connect the place to the scene and construct a sentence in another language. In the attention article, we saw how one position gathers information from others. Something remained unfinished: how to organize a network that uses that information to produce an output.
A Transformer is a way of organizing a neural network around attention. We can picture the whole journey before opening its parts: prepare the sentence, relate its pieces and use the result to construct a response. In the original design, two modules divide that work: the encoder processes the input, and the decoder consults that result to build the output.
The 2017 proposal dispensed with recurrent processing, which updates a state one position at a time. Attention could connect distant positions directly and compute many operations within a layer in parallel. Vaswani and colleagues evaluated the design mainly on translation (Vaswani et al., 2017 (opens in a new tab)). Our sentence lets us see how that organization works. We only need to follow what information is available at each point; the numbers can wait.
First, prepare the sentence
The computer starts by dividing the text into tokens, units that can be words, fragments or punctuation. It then assigns each one a list of numbers: its initial representation, or embedding. This is the step we explored in From words to vectors. Here we will use words as labels to follow the example, although a real tokenizer might split them differently.
So far, we have assigned numbers to each token. We also need to indicate their order in the sentence: “The fisherman found the cashier” and “The cashier found the fisherman” contain the same pieces, but who finds whom changes. Each token's initial representation needs to incorporate its place in the sequence so the model can use that difference.
The original Transformer adds a positional signal to each token's representation. The starting point therefore combines information about the token and its place. We can imagine cards laid on a table, each marked with a position. In the model, both the content and that mark are expressed through numbers; note 1 explains the detail.
The input is now ready. bank has a place at the end of our sentence, but its initial representation still needs to incorporate what is happening around it. That is where the encoder's work begins.
Relate the pieces and work on what they contribute
Think about what happens when we read. If we encounter bank alone, several interpretations remain open. Add a fisherman, his net and the act of repairing it, and the riverbank fits the scene. The model has a mechanism for combining information from those positions: the self-attention we met in the previous article.
In our example, the representation of bank can receive contributions from fisherman and net. Multi-head attention makes several mixtures of information in parallel and combines their results. Those mixtures are computed using parameters learned during training. The journey we imagine explains the mechanism; knowing which relationships a particular model uses requires examining it.
Attention has gathered clues from other positions into the representation of bank. A small neural network, called the feedforward network, then recombines its numbers. It can strengthen certain combinations of features and attenuate others, according to its learned parameters, preparing an update that later layers can use. Attention exchanges information between positions; the feedforward network works on what each one has gathered.
As these transformations are applied, residual connections keep a direct route for the incoming representation and add the update to it. At bank, this lets what was already there reach the sum alongside the new information. Normalization adjusts the scale of the numbers and helps keep training stable as the process repeats across many layers. Note 2 shows how both operations fit together and provides their calculations.
Together, these parts form a block. The next block receives the representations left by the previous one and works on them again. This lets bank incorporate new relationships using information that has already passed through other transformations. Each block has its own learned parameters, even when it repeats the same structure.
It is worth pausing here: we still have one representation for each input position. The encoder has prepared information that the decoder can consult. The translation is built in the next part of the journey (Vaswani et al., 2017 (opens in a new tab), section 3.1 (opens in a new tab)).
Build the translation, one step at a time
Suppose the output already says: “El pescador reparó su red junto a la…” This means “The fisherman repaired his net beside the…” The decoder needs to continue from there. It has two sources: what has already been written in Spanish and the information the encoder prepared from the English sentence.
First, it uses attention to relate the available parts of the translation. Here the causal mask from the previous article returns: each position can consult its own token and earlier ones, while later positions are blocked. When continuing from “la”, that token is also part of the context. The restriction concerns the output; the whole English sentence remains available.
Then comes cross-attention, which we also introduced when discussing translation. It connects the two sides of the journey. From the position holding “la”, the decoder uses what has been written so far to consult the encoder's information. In our scene, that consultation can receive a large contribution from bank, whose representation has already had opportunities to incorporate context from the fisherman and his net.
You can follow that meeting in the animation. The lines show how attention weight is distributed across positions in an invented example. What matters now is seeing how the original sentence and the unfinished translation connect. If you are curious about the calculation, open “See the numbers”.
Motion is paused. Use the step buttons to explore.
Original sentence · English
- The
- fisherman
- repaired
- his
- net
- beside
- the
- bank
Translation so far · Spanish
El pescador reparó su red junto a la
“The fisherman repaired his net beside the…”
Where we continue
- 1. Use what is already writtenAt “la”, the model has the translation written so far available as context. This helps determine what information to draw from the original sentence.
- 2. Consult the original sentenceThe original sentence contributes information to this position. In our example, the largest contribution comes from “bank”, whose representation already incorporates context.
- 3. Gather information to continueThe contributions are combined and pass through further transformations. Only afterward does the model assign probabilities to possible next tokens.
See the numbers
The query comes from the representation at “la” and is compared with encoder keys. These invented weights for one head add up to 1. Each multiplies a value vector; the sum is the contribution passed onward, not a probability of the next token. Source positions can be processed in parallel.
The0.02 × [0, 0]
fisherman0.10 × [1, 0]
repaired0.02 × [0, 0]
his0.02 × [0, 0]
net0.20 × [0, 2]
beside0.02 × [0, 0]
the0.02 × [0, 0]
bank0.60 × [2, 1]
0.60[2,1] + 0.20[0,2] + 0.10[1,0] + 0.10[0,0] = [1.3,1.0]
An invented example: a trained model may distribute attention differently.
The gathered information continues through the decoder's transformations. At the end, the model assigns probabilities to the tokens that could come next. In our example, orilla (“riverbank”) might receive a higher probability than banco (“bank” or “bench”) or casa (“house”) and be selected as the continuation. A word may require several tokens, and nothing guarantees that the choice will be correct.
There are two different operations along this stretch. Attention weights indicate how much the information from each position is weighted; its contribution to the mixture also depends on the values being combined. The final probabilities allow a token to be selected for the output. A thicker line toward bank is not a probability of writing orilla.
The selected token joins the text and the process continues. The decoder now has a slightly longer translation available to produce the next token, until a stopping condition is met. We can compute many internal operations in parallel and still generate the continuation step by step. Note 3 expands on that distinction.

The original sentence remains available while the translation grows. A conceptual illustration.
From translation to models that continue text
So far, we have used both parts of the original Transformer. Other architectures use related components with a different organization. Recognizing them helps connect our translation to other uses of language.
An encoder-only model, such as BERT, uses context on both sides of a position to build representations. These can support tasks such as classifying text or locating an answer within a passage (Devlin et al., 2019 (opens in a new tab)).
A decoder-only model, such as GPT-2, works with the available sequence to predict the next token. The input text and its continuation share a single path, with causal attention. In this design, there is no separate encoder to consult through cross-attention (Radford et al., 2019 (opens in a new tab)).
This second organization helps place the idea of continuing text that we encounter in conversational models. Training it to follow instructions and building a chat application involve additional decisions. For now, we can recognize the basic mechanism: use the available context to propose a continuation, then repeat the process.
Return to the fisherman
We began with an ambiguous word at the end of a sentence. Now we can follow its journey: positional information helps preserve order; attention incorporates context; layers transform what has been gathered; and the decoder combines the input with what has already been written to continue the translation.
Returning to orilla, we can now recognize how clues from the fisherman, the order of the sentence and the unfinished translation came together. That is the picture to take away: information being related and transformed, until it contributes to choosing what comes next.
The architecture provides those operations. Training adjusts the parameters that make their combination useful. Drawing the connections therefore still leaves an important question open: what will the model learn from the examples it receives? That will be the next stretch of the series.
To explore these parts for yourself, here are three short walkthroughs:
-
Follow a query in The Transformer, Drawn (opens in a new tab). In the tool I built, keep the example sentence and open “05 Scores”. Under “Query token”, select the first word and then the last: watch how the available positions change. The first can only attend to itself; the last can also attend to every earlier position. Move to “06 Softmax” and “07 Values” to follow scores becoming weights and weights producing a mixture of values. This is a simulation with invented numbers: it omits residual connections and normalization, and simplifies feedforward. You can inspect those choices in the code on GitHub (opens in a new tab).
-
Locate the whole journey with The Illustrated Transformer (opens in a new tab), by Alammar (2018). Start with the overall diagram separating encoder and decoder. Then find “The Residuals”: follow the arrows around each transformation and locate normalization. These are the parts my simulation omits. Continue to “The Decoder Side” and trace the connections from the encoder to the decoder's attention. You can relate them to our query from “la” to the English sentence; the drawing helps show where that meeting takes place within the architecture.
-
Try a continuation with Transformer Explainer (opens in a new tab). This resource runs GPT-2 in the browser (Cho et al., 2026 (opens in a new tab)). While the model loads, you can explore the sentences under “Examples”. Once it is ready, enter “The fisherman repaired his net beside the” and click “Generate”. Here you will see how it continues the sentence in English. Look at the candidates under “Probabilities” and the token that gets added. Then restore the same prefix and change only “Temperature”, keeping the sampling settings fixed: lowering it concentrates the distribution more on the highest-scoring candidates; raising it spreads it out. This exercise helps distinguish the available probabilities from the token that is eventually selected.
Notes: optional technical detail
These optional notes collect the mathematical details and variants. The main journey is complete without them.
-
Order and positional signals. In self-attention without positional information or a mask that distinguishes positions, permuting the input rows permutes the output rows in the same way: the operation is equivariant to that permutation, not invariant. Positional signals break that symmetry. The original Transformer's sines and cosines use different frequencies across coordinate pairs; the paper also compares learned positions (Vaswani et al., 2017 (opens in a new tab), section 3.5 (opens in a new tab)). A causal mask introduces an asymmetry: later positions can access more predecessors. Haviv and colleagues found that causal language models without explicit positional encodings can learn positional information; they propose that this access pattern helps explain it (Haviv et al., 2022 (opens in a new tab)). This is why the argument for unmasked attention cannot simply be carried over to the causal case. Return to positions.
-
The coordinates we normalize. For a vector with coordinates at one position, and . Then . The constant stabilizes division; and are learned parameters, shared across positions in that normalization. For a small example, use a scale of one and a shift of zero. With those settings and omitting for the arithmetic, , and all become : two coordinates already make the loss of scale visible. With a small positive , these results are approximate; adding the same constant to both input coordinates still leaves the result unchanged. Normalization does not preserve all vector information invertibly, and the learned transformation means its output need not have zero mean and unit variance (Ba et al., 2016 (opens in a new tab)).
The original encoder block can be written as and . Here, contains the input representations, is the intermediate result and is the output; each row represents one position. The feedforward network applies two learned transformations with a ReLU activation between them; its parameters are shared across positions within a block. The attention operation includes its output projection, and these expressions omit dropout (Vaswani et al., 2017 (opens in a new tab), sections 3.1–3.3 (opens in a new tab)).

The side paths show the residual connections. This diagram follows the original encoder's Post-LN order and omits dropout.
Other variants place normalization before attention and feedforward within the residual branches. This Pre-LN arrangement changes gradient behavior and training conditions (Xiong et al., 2020 (opens in a new tab)). The diagram shows one particular design. Return to the block.
-
The reference as training input. Feeding the decoder the correct prefix, rather than its own previous choices, is known as teacher forcing. The input is shifted so that each position predicts the next token; the mask prevents access to later positions. Errors in those predictions contribute to a loss, and their gradients adjust the parameters. During ordinary generation, the prefix grows with the tokens being selected, without access to a reference continuation (Vaswani et al., 2017 (opens in a new tab), section 3.1 (opens in a new tab)). In dense self-attention, comparing every query with every key produces a table with entries for positions. Doubling the sequence length quadruples that count; it is not a direct measure of latency or memory in every implementation. Return to the translation.
References
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention Is All You Need (opens in a new tab). In I. Guyon, U. von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, & R. Garnett (Eds.), Advances in Neural Information Processing Systems (Vol. 30, pp. 5998–6008). Curran Associates.
Alammar, J. (2018, June 27). The illustrated Transformer. Illustrated article (opens in a new tab)
Ba, J. L., Kiros, J. R., & Hinton, G. E. (2016). Layer normalization [Preprint]. arXiv. https://doi.org/10.48550/arXiv.1607.06450 (opens in a new tab)
Cho, A., Kim, G. C., Karpekov, A., Lee, S., Helbling, A., Hoover, B., Wang, Z. J., Kahng, M., & Chau, D. H. (2026). Transformer Explainer: Learning LLM Transformers with interactive visual explanation and experimentation. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (pp. 1–21). ACM. https://doi.org/10.1145/3772318.3791725 (opens in a new tab)
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional Transformers for language understanding. In J. Burstein, C. Doran, & T. Solorio (Eds.), Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) (pp. 4171–4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423 (opens in a new tab)
Haviv, A., Ram, O., Press, O., Izsak, P., & Levy, O. (2022). Transformer language models without positional encodings still learn positional information. In Y. Goldberg, Z. Kozareva, & Y. Zhang (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2022 (pp. 1382–1390). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.findings-emnlp.99 (opens in a new tab)
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners [Technical report]. OpenAI. Original report (opens in a new tab)
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., & Liu, T. (2020). On layer normalization in the Transformer architecture. In H. Daumé III & A. Singh (Eds.), Proceedings of the 37th International Conference on Machine Learning (Vol. 119, pp. 10524–10533). PMLR. Conference proceedings (opens in a new tab)