Key Features

Text-to-image generation is included in its unified task set.
Instruction-based image editing shares the multimodal backbone.
Image understanding and document or chart reasoning are evaluated.
The project reports video-understanding tasks and benchmark results.
3D understanding is included among the demonstrated capabilities.
Multiple discrete codebook levels represent each visual position.
Spatial positions are processed in parallel while codebook levels are decoded sequentially.
The repository supplies inference code and an Ascend-trained model checkpoint.

The design combines an encoder and task heads with a mixture-of-experts diffusion language-model backbone. Multi-codebook discrete visual representations encode each spatial position, with parallel processing across positions and sequential decoding across codebook levels. This connects visual generation and understanding within one modeling framework.


The project is aimed at multimodal researchers and developers exploring unified visual systems. The repository provides inference code and an Ascend-trained checkpoint. Its NVIDIA-trained checkpoint, full training code, and technical report are listed as later releases, while the project metrics are described as preliminary.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!