Key Features

The model retains the pretrained Qwen3.5-4B vision-language architecture.
It jointly predicts 3D detections, semantic occupancy, and bird-eye-view map segmentation.
The unchanged vision-language backbone supports free-form scene question answering.
A flow-matching Planning Expert predicts future ego trajectories.
The model includes planner-sft and a reward-optimized planner-rl variant.
The SFT planner supports direct and reasoning modes; RL favors reasoning mode.
Results include open-loop, pseudo-closed-loop, and closed-loop driving benchmarks.
The ModelScope model card lists Apache 2.0.

The system retains the pretrained Qwen3.5-4B architecture and attaches an external bird-eye-view perception head and a flow-matching Planning Expert. Staged training combines driving supervision with general visual-language data. Separate SFT and reward-optimized planners support different planning modes and evaluation settings.


The release is useful for autonomous-driving research involving interpretable perception and language-conditioned planning. ModelScope provides weights and configurations, with code and demo data in the companion repository. It is a research model requiring an appropriate evaluation and integration environment, not an independently validated vehicle-driving product.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!