What Is Attention? Transformer Models

Editorial archive of Creangel. The references, figures and conditions correspond to the original publication.

Introduction to the Attention Mechanism in Transformers

What is Attention?

Attention is like when you read a phrase and your brain decides which words are most important to understand the meaning.

Introduction to the Attention Mechanism in Transformers

Introduction to the Attention Mechanism in Transformers

Based on.

Scaled Dot-Product Attention

It's a mathematical formula that works like this:

  1. Compare: Take a word (query) and compare it with all other words (keys).
  2. Calculate Weights: Gives More Importance to the Most Related Words
  3. Combine: Mix all information according to those weights

Formula: Attention(Q,K,V) = softmax(QK^T/√dk)V

Multi-Head Attention

Instead of using a single “head” of attention, use several (8 in this case):

  • Each head focuses on different aspects of the relationships between words
  • It's like having 8 different perspectives from the same text.
  • At the end all perspectives are combined

Three Uses in Transformer:

  1. Encoder-Decoder: The decoder looks at the entire input
  2. Encoder self-attention: Each word looks at all other words in its layer.
  3. Self-attention in Decoder: Each word only looks at the previous ones (can’t “cheat” seeing the future)

In short: It is a mechanism that helps the model understand which parts of the text are most relevant to each word.

In the formula Attention(Q,K,V) = softmax(QK^T/√dk)V, T means transpose.

What is the transposition?

The transposition of a matrix means turning the matrix by swapping rows into columns.

Example:

Original matrix K: Transposed K^T:
[1 2 3] [1 4]
[4 5 6] [2 5]
[3 6]

Why is K^T used?

To be able to multiply matrices between Q and K:

  • Q has dimensions: (number of queries × dk)
  • K has dimensions: (number of keys × dk)
  • K^T has dimensions: (dk × number of keys)

Then:

  • Q × K^T = (number of queries × dk) × (dk × number of keys)
  • Result = (number of queries × number of keys)

This gives us a matrix where each element represents how compatible each query is with each key.

In Python code:

python

scores = torch.matmul(query, key.transpose(-2, -1))

The .transpose(-2, -1) is equivalent to setting the T (transposed).

dk is the dimension of keys and queries.

Simple Definition

dk = the number of elements (features) in each key and query vector.

In the Context of Transformer

According to the document:

  • dk = dv = dmodel/h = 64

Where:

  • dmodel = total model size (e.g. 512)
  • h = number of attention heads (8 in this case)
  • dk = 512 ÷ 8 = 64

Why is it important?

It is used in two key locations:

  1. In the division by √dk:

scores = QK^T / √dk

  • If dk = 64, then √dk = 8
  • This normalizes scores to keep them from being too large.
  1. In the dimensions of the matrices:
  2. Each query vector (Q) has 64 elements
  3. Each key vector (K) has 64 elements
  4. Each vector of value (V) also has 64 elements

Why Divide by √dk?

As the paper explains: when dk is large, the dot products QK^T become very large, causing the softmax function to have very small gradients. Dividing by √dk stabilizes training.

Example: If dk = 64, we divide by √64 = 8 to keep values within a reasonable range.

Let's talk about your information.

Tell us what you need to find, analyze or organize. Discover IFINDIT with the Creangel team.

Image