1. Introduction
2. Limitations of Transformers
3. Titans: Architectures Inspired by Human Memory
4. Advantages of Titans over Transformers
5. Three Ways to Incorporate Memory
6. References
This note aims to present the architecture introduced in Titans: Learning to Memorize at Test Time. It compares Titans to traditional Transformers, highlights its advantages, and describes the three proposed strategies to integrate memory.
One problem with Transformers is the quadratic complexity of the attention mechanism, causing performance and viability to decrease rapidly as the input sequence length grows.
For example, in tasks where contextual dependencies span very long sequences—such as large documents, biological data, or time series—standard Transformers cannot efficiently handle these sequences without using truncations or simplifications. These limitations can negatively affect result quality.
The following chart compares three complexities: quadratic, linear, and logarithmic.
Source: Own elaboration
This chart clearly shows how quadratic complexity limits the use of standard Transformers for long sequences.
To address these issues, in the paper Titans: Learning to Memorize at Test Time, Google researchers propose a new architecture, Titans, that overcomes these limitations. This architecture is based on the structure of human memory to integrate short-term memory, long-term memory, and persistent memory.
To this end, Titans includes a neural memory module that combines three main types:
- Short-term memory that captures local dependencies, similar to the attention mechanism of Transformers.
- Long-term memory that summarizes and stores important historical relationships.
- Persistent memory that encapsulates the general knowledge acquired during training.
Memory operates dynamically, adjusting based on the relevance of the data. Titans uses a “surprise” metric to identify which information should be stored in long-term memory, emulating the human ability to recall significant events. It also introduces a forgetting mechanism that discards irrelevant information, optimizing the ability to handle new data.
While Transformers must split long sequences into chunks due to their quadratic complexity, Titans can handle sequences of millions of tokens with linear or sub-quadratic complexity, thanks to its long-term memory module preserving global relationships without processing all possible token combinations.
In addition, Titans uses optimized mini-batch operations, significantly reducing memory and computation costs. This design makes it a much more efficient solution for tasks requiring extensive sequence analysis.
The advantages of Titans are remarkable. It can process contexts of up to 2 million tokens without loss of accuracy, making it exceptionally efficient. Its ability to dynamically update memory during inference allows it to adapt to new data without retraining.
These features establish Titans as a powerful, advanced solution to the traditional limitations of Transformers.
Titans incorporates memory in three distinct variants to efficiently address the limitations of traditional Transformers. These variants focus on combining short-term memory, long-term memory, and persistent memory, adapting to different needs in deep learning tasks.
1. Memory as Context (MAC)
Source: Titans: Learning to Memorize at Test Time
In this variant, long-term memory is treated as an additional context that complements the current input.
2. Memory as Gate (MAG)
This variant uses long-term memory as a separate branch that interacts with the model’s main flow through a non-linear gating mechanism.
Source: Titans: Learning to Memorize at Test Time
3. Memory as Layer (MAL)
Here, long-term memory is integrated as an autonomous layer within the architecture, acting like a recurrent or linear transformer in a deep model.
Source: Titans: Learning to Memorize at Test Time
Comparison and Variant Selection
| Variant | Main Benefits | Ideal Scenarios |
|---|---|---|
| MAC | Efficiently combines historical memory and current context. | Tasks with extensive dependencies and long contexts. |
| MAG | Flexibility and efficiency via dynamic gating. | Applications where local and global dependencies are equally important. |
| MAL | Simple and straightforward; suitable for sequential structures. | Hybrid models or tasks with clear patterns. |
Your reading notebook
The note is saved only in this browser.
This archive note retains its original publication context.
View original archive file ↗