Artificial intelligence

Robots, PacMan and the laws of learning

Reinforcement learning, interaction with the environment and agent adaptation.

5 min readFeb 2025Archive note
English edition on the blog ↗

The image shows a robot playing Pac-Man

Introduction

Let’s suppose we are learning a new skill, such as playing tennis. We might ask ourselves different questions:

  • Would it be better to train directly on an outdoor court, with wind, noise, and changing light conditions?
  • Or would it be more effective to start in a controlled environment, free from distractions?

Our intuition tells us that training in conditions as similar as possible to real competition should yield the best results. But, what if that weren’t the case?

Isaac Asimov revolutionized science fiction with his concept of positronic robots, endowed with advanced intelligence and regulated by his famous Three Laws of Robotics. These laws, designed to ensure the safety and functionality of robots, stated that:

  1. First Law: A robot may not injure a human being or, through inaction, allow a human being to come to harm.
  2. Second Law: A robot must obey the orders given by human beings, except where such orders would conflict with the First Law.
  3. Third Law: A robot must protect its own existence as long as such protection does not conflict with the First or Second Law.

How would the laws of efficient learning for a robot be formulated?

A Surprising Discovery in Robot Learning

A recent study on reinforcement learning, titled The Indoor-Training Effect: Unexpected Gains from Distribution Shifts in the Transition Function, aimed at improving robot training, has found a surprising result: in some cases, training in a cleaner and more structured environment can lead to better performance in difficult conditions.

This finding, called the Indoor-Training Effect (ITE), challenges the conventional idea that the best training is the one that exactly resembles the environment where the knowledge will be applied.

The Discovery in Artificial Intelligence

This research analyzed how AI agents learn in different environments through classic ATARI games (Pac-Man, Pong, and Breakout). Two types of agents were compared:

  • Learnability Agent (Lδ): trained and tested in the same noisy environment.
  • Generalization Agent (GT): trained in a noise-free environment and tested in a noisy environment.

Against all expectations, the Generalization Agent outperformed the Learnability Agent in many tests. In other words, training in a cleaner environment allowed agents to perform better in more complex environments. The Indoor-Training Effect suggests that starting in a controlled environment can enhance the ability to adapt to challenging scenarios, rather than simply getting accustomed to the noise of the final environment.

Phase 2: Games Used in the Study

This research analyzed how AI agents learn through classic ATARI games: Pac-Man, Pong, and Breakout.

Games and Applied Rules

Researchers selected these games because of their fast-paced dynamics and the need for strategic planning, making them ideal for evaluating the generalization capabilities of AI agents. For example, in Pac-Man, where the goal is to eat all the dots without being caught by ghosts, noise modifications involved changes in the speed and trajectory of the ghosts.

(Image here) The following image, obtained from the study, shows the training process in Pac-Man

Source: The Indoor-Training Effect

The diagram below illustrates the three phases used to structure the experiment:

  • Environment Creation
  • Training
  • Evaluation

These phases were used to test this methodology.

Source: Own elaboration

Parallels with Human Learning

This phenomenon in artificial intelligence mirrors a pattern observed in the neuroscience of human learning. Over the course of evolution, organisms have developed mechanisms to learn in a structured manner first, and then generalize to more complex environments.

For example, a beginner tennis player first practices on an indoor court, with a machine that launches balls predictably. Only after mastering the technique in that setting does the player face an opponent, deal with wind changes, and experience the pressure of a real match. This method helps the brain solidify movement patterns before confronting the variability of the real world.

Final Reflection and the Laws of Learning for Robots

In the development of robots and artificial intelligence, it has long been assumed that training in environments that exactly mimic real-world conditions is the best strategy for optimal performance. However, the Indoor-Training Effect (ITE) challenges this belief, suggesting that training first in a clean and structured environment can lead to better performance in difficult conditions.

Likewise, in robotics, the ITE paradoxically suggests that training first in a clean and structured environment can improve performance in challenging conditions, drawing a parallel with human education.

In the case of robots, one might imagine three laws of learning, analogous to Asimov’s Laws of Robotics:

  1. First Law of Learning: A robot must first learn in a structured and controlled environment before facing the complexity of the real world, unless doing so hinders its ability to adapt.
  2. Second Law of Learning: A robot must be gradually exposed to uncertainty and environmental noise, provided that this does not compromise the knowledge acquired during the initial training phase.
  3. Third Law of Learning: A robot must be capable of adapting to unexpected conditions, as long as this does not involve unlearning what was learned in the initial phase or generating systematic errors in its performance.

One could even imagine, just as Asimov did, a Zeroth Law of Learning for robots:

  • 0️⃣ Zeroth Law of Learning: A robot’s learning must optimize its performance in the real world, even if this requires modifying the previous rules.

References

The Indoor-Training Effect: Unexpected Gains from Distribution Shifts in the Transition Function

Make It Stick: The Science of Successful Learning by Brown, Roediger & McDaniel (2014)

View Reference

Your reading notebook

The note is saved only in this browser.

This archive note retains its original publication context.

View original archive file ↗