This is the model's memory: a table with 1,048,576 entries, each a vector of 384 numbers. Three memory layers (layers 3, 7 and 11) share it. For every token, each of them picks the 4 × 32 = 128 best-matching entries by product-key search and mixes their vectors. The numbers alone don't tell you much. It gets interesting when you see when an entry is read.
This is B-1M from the early runs. The B-16M model in the README works the same way, with a table 16 times as big.
Click a word. Below it you see, for each memory layer, the three entries with the highest weight for exactly that token. The texts are from the validation set, the model never trained on them.
The table as a 1024 × 1024 grid. The row is the number of the first half-key, the column the number of the second. An entry is picked when both halves of the query match it well. The colour shows how often the entry was read during training tokens. Solid lines are very popular half-keys. Click the map to pick an entry, the box next to it shows its neighbourhood.
Picked by hand from the sample. They show when an entry is read, which isn't the same as knowing what it does there.