Show original animation

Image Representing a Digital Environment with Two Versions of AI Competing on a Chessboard
1. From LLMs to LRMs
Artificial intelligence has evolved from simple predictive models to systems capable of reasoning and improving their own performance without human intervention. This evolution has led to the emergence of Large Reasoning Models (LRMs), which go beyond text generation to explore multiple solution paths, correct errors, and develop structured thinking. This can be observed in models such as OpenAI’s latest releases, Google’s Gemini Thinking Model, and Deepseek R1. In this article, we explore how techniques like RLSP (Reinforcement Learning via Self-Play) allow AI to learn to reason by itself, getting ever closer to human-like thinking.
What is the basis of this transformation?
This transformation is based on the ability of reasoning models to conduct deliberate searches and execute logical reasoning chains during inference. Unlike traditional LLMs, whose primary goal is to generate coherent text, LRMs can explore different solution paths, verify their responses, and iteratively correct their mistakes. This approach not only improves accuracy but also allows models to tackle more complex problems with structured reasoning.
For example, while a conventional language model may solve a mathematical equation by applying a known formula, a reasoning model can analyze the problem from multiple perspectives, evaluate different solution methods, and correct its conclusions if it detects inconsistencies in its reasoning.
This paper aims to explore this question by analyzing the proposed RLSP framework and its contribution to the development of models with advanced reasoning capabilities.
2. Games and Artificial Intelligence
The core of this study is the RLSP (Reinforcement Learning via Self-Play) framework, a post-training reinforcement learning method designed to improve the reasoning ability of language models. Unlike traditional approaches such as supervised learning or reinforcement learning with human feedback (RLHF), RLSP introduces a mechanism that allows models to learn autonomously by exploring different problem-solving strategies without constant human intervention.
The RLSP framework is based on three essential components that work together to promote deeper and more structured reasoning.
a) Supervised Fine-Tuning (SFT)
The first step in RLSP involves Supervised Fine-Tuning (SFT), where the model is trained with examples of structured reasoning. These examples can be provided by humans or synthetically generated using techniques such as Chain of Thought (CoT), which break down complex problems into logical sequences of intermediate steps.
b) Exploration Reward
The second component introduces an exploration reward, a key mechanism to encourage the model to seek innovative solutions beyond pre-existing patterns in its training data.
c) Reinforcement Learning with PPO
The third component of RLSP is reinforcement learning using the PPO (Proximal Policy Optimization) algorithm, a widely used technique for optimizing AI models without causing abrupt behavioral changes. To ensure the model does not deviate too far from its initial behavior, a penalty based on KL divergence (Kullback-Leibler) is applied. This prevents the model from generating arbitrarily long responses without useful content (reward hacking), maintaining a balance between exploration and adherence to prior knowledge.
This process is illustrated in the following diagram:
(Diagram source: Author’s creation)
A fundamental aspect of the RLSP framework is its difference from RLHF (Reinforcement Learning from Human Feedback), a method commonly used to improve language models through direct human feedback.
Comparison: RLHF vs RLSP
| Characteristic | RLHF | RLSP |
|---|---|---|
| Feedback Source | Humans rate model responses. | The model learns by exploring multiple solutions on its own. |
| Learning Method | Based on explicit human rewards. | Based on an automatic verifier and exploration rewards. |
| Level of Autonomy | Depends on constant human intervention. | Allows the model to self-train without external supervision. |
| Scalability | Limited by the number of human evaluations available. | Scalable to any domain without requiring large amounts of human annotations. |
(Table source: Author’s creation)
While RLHF is useful for aligning the model with human preferences, its dependence on human evaluators limits its scalability. In contrast, RLSP allows the model to learn autonomously, optimizing its reasoning ability through self-assessment and exploration.
3. Emerging Behaviors in Models
One of the most significant findings in the implementation of RLSP is the emergence of unexpected behaviors in language models. These behaviors, which were not explicitly programmed, arise as a result of the optimization process based on exploration and self-evaluation.
Unlike models trained through supervised learning or RLHF, models optimized with RLSP demonstrate the ability to adjust their reasoning in real-time, making them more flexible and effective in solving complex problems.
Training with RLSP induces behaviors that resemble human thinking, especially in tasks requiring multi-step reasoning. Some of the most notable behaviors include:
- Backtracking (Revisiting and Correcting): When the model detects a potential inconsistency in its response, it can retrace its steps, review its calculations, and modify the solution based on new information.
- Exploration of Multiple Paths: Instead of following a single problem-solving method, the model evaluates different strategies before making a final decision, allowing it to find optimal solutions in complex cases.
- Self-Verification: Before providing a final response, the model checks the validity of its solution using techniques such as numerical verification or comparison with alternative methods.
Conclusion
As autonomous reasoning models continue to evolve, we are approaching a point where AI could surpass human capability in certain cognitive tasks. This raises fundamental questions about its role in decision-making, the ethics of its development, and its impact on society. Self-learning AI is not just a technological advancement; it represents a paradigm shift in our relationship with machines.
References
On the Emergence of Thinking in LLMs I: Searching for the Right Intuition
Your reading notebook
The note is saved only in this browser.
This archive note retains its original publication context.
View original archive file ↗