Transformers and Attention

The transformer is the architecture behind modern language models. Its key idea is attention: when the model reads a word, it can look at the other words in the same text and decide which ones matter for the meaning.

Why attention replaced older sequence models

Older recurrent networks read a sentence one word at a time and compressed everything into a single running state. Long sentences lost early details. Attention lets every token compare itself with every other token in the window, so a pronoun can connect to the right noun even when they are far apart.

Self-attention in one pass

  1. Each token is turned into a vector (an embedding).
  2. From that vector the layer builds a query, a key, and a value.
  3. The query is compared with every key. Higher scores mean those tokens are more relevant.
  4. The values are mixed according to those scores. The token’s new representation carries the context it needs.

A transformer block repeats this, often with several attention heads looking at different relationships at once, then a small feed-forward network. Stack many blocks and you get a deep language model.

Encoder, decoder, and decoder-only

DesignTypical jobExample use
EncoderRead a full text and build a representationSearch embeddings, classification
Encoder-decoderRead one text and write anotherTranslation, summarization
Decoder-onlyGenerate the next tokenChat models such as GPT-style LLMs

A practical example

In the sentence “The loan that the branch approved yesterday was booked today,” attention lets “was booked” attach to “loan,” not to “branch” or “yesterday.” That link is what makes the rest of the sentence interpretable.

What to remember

  • Attention is a weighted look at other tokens in the same context.
  • Transformers stack attention so meaning can depend on the whole passage, not only the previous word.
  • Most chat LLMs are decoder-only transformers trained to predict the next token.