Tools, Attention visualizerAll tools

What the model is looking at

Method

A toy multi-head self-attention layer. Enter a sentence, choose a head count and model dimension, and read the matrix. The math is real, the weights are randomly initialised.

I Configuration

0 tokens · 4 heads · d=64

parallel attention mechanisms

richer representations

II Tokens

embedding dims · 64

III Attention heads

active · head 1/4

IV How self-attention works

III stages

  • Stage · i

    Query, key, value

    Each token is transformed into three vectors, what to look for, what can be offered, and the content itself.

  • Stage · ii

    Attention scores

    Dot products between query and key vectors, then softmax, producing the attention weights.

  • Stage · iii

    Weighted sum

    Multiply attention weights with value vectors. That yields the final output for each token.

Multi-head attention transformers run several of these in parallel. Each head is free to specialise on a different relationship: syntax, semantics, a hunch about coreference. Together, they make the model rich.

V Try these

4 sentences

Attention Visualizer, Shagun Mistry