Artificial intelligence

Huginn: latent reasoning

An exploration of recurrent reasoning and internal computation beyond chains of words.

13 min readApr 2025Archive note
English edition on the blog ↗

Today's language models achieve their performance thanks to the intensive use of tokens: each step of reasoning is translated into text. This strategy, known as chain-of-thought reasoning, improves results, but imposes a structural limit: it forces all thought to pass through the filter of language.

In this article we analyze "Huginn", an innovative architecture based on recurrent latent reasoning, which allows models to internally process multiple steps before uttering a word. Huginn introduces a new axis of scalability: more depth of thought, not more parameters. An approach that challenges the current paradigm and opens up new possibilities for the future of artificial intelligence.

Cartography of Latent Thought

Show original animationAnimated visualization of latent space
Open animation ↗

The image represents a video made by Sora to represent the latent space. Although Sora does not use a latent space, but a pixel space for video generation, both spaces have much in common. The pixel space, like this image, represents the actual color of each point, while the latent space represents abstract information: contours, textures, semantics.

1. Introduction: What does it mean to "reason without words"?

For years, advances in language models such as GPT or ChatGPT focused on teaching them to predict words more and more accurately. These models became adept at generating fluent text, answering questions, and even solving complex problems, all based on one core skill: anticipating the next word in a sequence.

But in the midst of this evolution, a question began to take shape among researchers: do models really think while they write? Or are they just stringing together words that "sound good"? And if they can think... why force them to translate every step of their reasoning into words?

This is the disruptive idea behind the work we are going to explore: that language models could improve, and much so, if they were not forced to "talk" while reasoning. That perhaps, like humans, they could have a more abstract, richer, more flexible inner thought, and only turn it into words at the end, when they are already clear about what they want to say.

The current limitation is that these models need to express their ideas through tokens, basic units of text (such as words, fragments or characters). Every time they take a step in a reasoning, they must issue a new token. Not only does this consume time and resources, but it also forces the model to impoverish their thoughts, translating them into linear, simplified language.

And that's where a new proposal appears: to allow the model to reason "silently", in what scientists call latent space, a mathematical space where information is represented internally, in high-dimensional vectors, without going through words.

This is the hypothesis that opens the door to a new generation of models: more intelligent, more efficient, and above all, less dependent on language to think.

2. The current problem: thinking like humans, but talking like robots

To understand why this new line of research is so important, we first need to look at how the most advanced language models work today. Although they seem intelligent and versatile, there is a technical detail that conditions their entire way of "thinking": they must do it through tokens.

A token is nothing more than a basic unit of text: it can be a whole word, a syllable, a letter, or even a punctuation mark. Every time a model responds, it does so by generating a sequence of tokens, one by one. And each of those tokens isn't just a random word — it represents a decision the model made, based on all of the previous tokens.

Now, when a model has to solve a complex problem, such as a long sum, a logical question, or a scientific explanation, the most effective strategy discovered so far is the so-called chain-of-thought. This technique has the model write down their reasoning step by step, as if they were explaining aloud how they arrive at the answer.

The problem is that forcing the model to express each step in words consumes a lot of memory, processing, and context. It's like a person solving a math problem by saying out loud, "Now I'm multiplying by two. Now I subtract five. Now I compare the results...". If we think about it, this effort is not necessary to elaborate an answer, because we do the process internally, in silence.

The same happens with these models: every time they generate a token they are wasting part of their computational capacity on expressing themselves, instead of focusing on reasoning. In addition, by having to convert their ideas into text, they lose precision. Not everything that is thought can be said clearly, and even less so in a format limited by a sequence of words.

In other words: current models are designed to speak well, but not necessarily to think well. And that puts a ceiling on their capacity.

The big question that arises is: what if we could allow them to think like we humans do? That is, to process, compare, imagine and reason without having to put everything into words.

That's where the idea of reasoning in latent space, without tokens in between, comes in. As if the model could "turn a problem over" internally, deepen its analysis, try alternative paths… and only then, when it is ready, generate the final answer.

Latent token map in the initial step of reasoning

Show original animationInitial latent token map
Open animation ↗

Source: Authors' elaboration based on the Scaling up Test-Time Compute with Latent Reasoning repository

The image above represents a visualization of the recurrent latent reasoning process of an advanced language model. Each point corresponds to the internal state of a token of the input in the first reasoning step. The horizontal axes (PCA1 and PCA2) represent a reduced projection of the high-dimensional latent space using Principal Component Analysis (PCA), while the vertical axis indicates the position of the token in the original sequence of text.

This image is step 0 of a full animation (captured in the associated GIF), showing how the internal state of each token evolves with each iteration of the model's recurrent block. Throughout this sequence, the model applies a reasoning function multiple times on its latent representation, allowing it to refine its understanding before producing a verbal output.

At the initial instant, the states are grouped vertically, reflecting that no internal reasoning has yet occurred: all tokens have been barely embedded in the latent space, but they have not interacted or refined their meaning.

However, as you progress through the iterations (as seen in the animation), each token begins to chart a unique trajectory in the latent space. Some of these trajectories:

  • Extend widely, indicating tokens that require more reasoning steps (e.g., key concepts such as “dozen” or “weeks”).
  • Remain close to their starting point, reflecting trivial or functional tokens such as prepositions or determiners.
  • Form loops, spirals, or straight lines, evidencing internal reasoning patterns such as orbits (iteration), sliders (progressive counting), or fixed-point convergence (stabilized understanding).

This visualization reveals a critical aspect of the new paradigm of reasoning in language models: thought occurs silently, in the latent space, without the need for strings of intermediate words. Instead of explicitly reasoning with text as in classic chain-of-thought approaches, the model internally refines its understanding before issuing an answer. Each step transforms the latent vectors of the tokens, allowing progressive reasoning without verbalization.

3. The Huginn model: a mind that reasons in silence

The model is designed to think deeply without necessarily putting each step into words.

The key lies in a simple but powerful idea: reuse its own inner layers to reason longer, without generating text until the end. This is what the authors call recurrent depth.

The model receives input, transforms it into an internal vector (like a kind of abstract thought), and then repeats the reasoning process on that vector as many times as it needs. Only when it feels that it has reached a conclusion does it translate that vector into text.

This has two major consequences:

  1. The model can reason deeper with the same size. Although Huginn has 3.5 billion parameters (a modest amount compared to today's large models), it can simulate being much deeper simply by repeating its reasoning steps. In computational terms, this is like having a model of 50 billion parameters... without having to train one.
  2. Reasoning is kept in latent space, that is, in an internal format that does not need to be translated into words at every step. It's as if the model could think quietly, abstractly, without wasting time or precision in converting every thought into language.

This represents a paradigm shift: we went from models that generate more tokens to think better, to models that use more internal computing without the need to talk more.

4. How do you train a model that thinks without speaking?

In classic models, the training process basically consists of showing it millions of pieces of text and asking it to predict the next token. Every time it makes a mistake, it adjusts its parameters a little. This cycle repeats itself billions of times.

But with Huginn the game changes: now you have to teach it not only to predict words, but to use its own inner layers as an iterative reasoning tool. It's like teaching it not only to speak, but to reflect before speaking.

The key new feature is that Huginn has a recurrent block: a set of layers that it can apply over and over again. But for this to work, you have to train it with different "rhythms of thought". That is, during training it is taught to solve tasks by doing from 1 to 64 internal repetitions of that block.

Each sequence of text shown to it during training goes through a different number of iterations, chosen at random. Thus, the model learns to improve its responses when it is given more time to think, and also to solve things quickly when it does not require as much processing.

This is called variable recurrence training, and it's what allows Huginn to dynamically adapt during inference (when it's already in use).

Training a recurrent model is not easy. Repeating the same operation many times can cause the information to collapse or explode: either it loses its meaning (everything flattens out), or it becomes chaotic.

To avoid that, the researchers used several clever techniques:

  • Initial random state injection: Each sequence starts with a distinct internal vector, which helps prevent the model from falling into rigid patterns.
  • Controlled feedback: During training, errors are not calculated for all iterations, but only for the last ones. This reduces memory consumption and prevents the model from being overwhelmed trying to adjust each internal step.
  • Specific norms: Carefully placed normalization layers (such as RMSNorm) are applied to maintain the numerical stability of the model while reasoning.
  • Continuous and modular training: the model is structured in three blocks (prelude, recurrent, coda) that can be tuned independently.

All of this is trained on a supercomputer with thousands of GPUs, in batches of up to 12 hours per segment. The result: a model with only 3.5 billion parameters that learned to behave as if it had many more, simply because it can "think deeper."

5. Amazing results: smaller, more efficient, smarter

Huginn was trained with 3.5 billion parameters. To put it in context, it's an intermediate size, much smaller than large commercial models like GPT-3 (175B) or PaLM (540B), and even smaller than many modern open source models like Mistral (7B) or Mixtral (12.9B).

And yet, Huginn achieves comparable—and sometimes superior—performance on several tasks that require complex reasoning.

How? Thanks to its ability to scale reasoning, not size. That is, it can internally repeat its own thought processes, rather than relying on more fixed layers or more training tokens.

In the most common evaluation tests for language models, Huginn showed progressive improvements as it was allowed to use more internal iterations in its recurrent block (i.e., think more).

Most impressively, performance is continuously improved with more recursion, proving that its recurrent architecture is not a one-time gimmick, but a robust avenue for scaling intelligence.

6. Emergent behaviors: when the model learns to think on its own

How Huginn thinks: three styles of internal reasoning

Huginn doesn't just solve tasks. It learns to reason in different ways depending on the challenge. This infographic summarizes the three emerging patterns that the researchers identified:

  • Convergence for simple tasks.
  • Orbits to manipulate abstract or spatial concepts.
  • Sliders for complex or moral evaluations.

Emerging patterns of latent thinking in AI

One of the most fascinating discoveries of this model is that not all thoughts follow the same internal path. Depending on the task, Huginn reasons in different ways, and this can be seen directly in its latent space.

The image below shows how the state of certain tokens is internally transformed as the model reasons silently. Each trajectory corresponds to a specific word as it goes through multiple iterations of the recurrent block. What is remarkable is that the model adopts different patterns of reasoning spontaneously, without anyone having taught it:

  • Rapid convergence: some words stabilize quickly, as if the model "knows" that they do not need further reflection.
  • Latent orbits: In spatial or numerical tasks, internal vectors rotate in circular patterns, as if Huginn were manipulating complex ideas in its mind.
  • Internal sliders: for abstract, moral or ambiguous reasoning, the model advances in a single and progressive direction, as if adjusting an internal scale of evaluation.

These emergent behaviors point to a powerful idea: Huginn doesn't just solve tasks, it learns to manage its own way of thinking. This brings it closer to a more flexible, adaptive, and, in a sense, strategic type of reasoning.

Latent trajectories of trivial tokens during recurrent reasoning

Show original animationEmergent patterns of thinking
Open animation ↗

Source: Authors' elaboration based on the Scaling up Test-Time Compute with Latent Reasoning repository

The visualization above shows the evolution of the latent state of selected tokens throughout the model's internal reasoning process. Each row corresponds to a specific token in the text (for example, "3", "many", "?"), and each column shows its projected path in different planes of latent space reduced by principal component analysis (PCA).

Each subgraph represents two main components of latent space. In them, each colored dot corresponds to the latent state of the token in an iteration of reasoning. The color scale indicates the time point within the process, where the darkest colors represent the first iterations and the lightest represent the last. The red cross indicates the center of mass of the trajectory, functioning as a reference to observe the movement of the token.

In this particular image, all tokens show extremely short trajectories, practically reduced to a single point. This indicates that there was no meaningful reasoning for these tokens. The latent state of each one did not change substantially over the course of the iterations, or it did so imperceptibly. In cognitive terms, the model quickly identified that these tokens did not require additional reflection.

This is explained by the nature of the tokens selected, which include punctuation marks (such as "," or "."), functional words ("many"), trivial numbers ("3"), and structural symbols such as the question mark ("?"). These elements tend to have fixed or highly predictable semantics, so the model understands them correctly from the initial embedding stage. As a result, there is no need to "think more" about them. The latent reasoning quickly stops and the trajectory converges in the first iterations.

This behavior clearly exemplifies how the model allocates more computational resources to the elements that require it, and avoids spending processing on those it already understands with certainty. It is a form of adaptive reasoning, where reflection is distributed efficiently, simulating a type of internal cognitive attention.

7. Conclusion: The future of artificial intelligence might not be so talkative

This model is not notable for its size or amount of data, but for its ability to reason silently, drill down when you need to, and stay efficient when you don't. Like humans, it manages its cognitive effort according to the problem. Sometimes it resolves quickly. Other times, it takes its time.

It is a more sober, more efficient and, perhaps, closer to how real thinking works, both in us and in the machines we are building.


References

Your reading notebook

The note is saved only in this browser.

This archive note retains its original publication context.

View original archive file ↗