Artificial intelligence

Janus-Pro and multimodal artificial intelligence

Image generation and new tools in China’s AI ecosystem.

2 min readFeb 2025Archive note
English edition on the blog ↗

Image generated by Janus-Pro

🌐 Change Language

The aim of this note is to analyze the new image generator model, its architecture, comparing it with diffusion models.

1.1 What is Janus-Pro?

Janus-Pro is an advanced multimodal model developed by DeepSeek-AI, designed to integrate and optimize the capabilities of image and text comprehension and generation within a single framework. It is a direct evolution of its predecessor, Janus, with significant improvements in performance, scalability, and stability.

2.1 Decoupled Architecture

Janus-Pro addresses one of the main challenges in multimodal models: the interaction between visual comprehension and generation tasks. To solve this, it introduces:

  • Comprehension Encoder: Extracts semantic information from images.
  • Generation Encoder: Converts images into discrete tokens (visual IDs).

Both feed a unified autoregressive transformer that integrates and aligns text and images into a single representation.

2.2 Generation Process

The generation process is optimized to translate text into high-quality images:

  1. Converting text to tokens
  2. Translating them into visual features

Advantages: semantic accuracy, stability, versatility.

The figure shows how Janus-Pro handles multimodal comprehension (left, blue) and image generation (right, yellow), both connected to a unified autoregressive transformer.

Diffusion models (e.g., Stable Diffusion, DALL-E) gradually remove noise from an image. Janus-Pro, however, uses an autoregressive transformer that generates images token by token, reducing the need for multiple iterations.

Comparison Table: Diffusion vs. Janus-Pro

Conclusion: While diffusion is powerful for high-resolution images, Janus-Pro's autoregressive approach is lighter and better for multimodal tasks.

Online demonstration of Janus-Pro on Hugging Face Spaces, from which an image was generated using a prompt created by DeepSeek-R1.

Example prompt for a night sky with hundreds of floating red and golden lanterns.

References

Your reading notebook

The note is saved only in this browser.

This archive note retains its original publication context.

View original archive file ↗