Qwen introduces Qwen-Drive-1.0, a vision-language foundation model for autonomous driving
· By Antonio Sedino, CTRO · Published by Reinventy Solutions Corp.
New model unifies 3D perception and visual question answering at pretraining stage, extending to motion planning while keeping pretrained VLM architecture untouched

Researchers from Qwen have introduced Qwen-Drive-1.0, described as the first vision-language foundation model for autonomous driving. The model unifies 3D perception and visual question answering at the pretraining stage and further extends to motion planning, while keeping the pretrained VLM architecture entirely untouched. The model is built on the natively multimodal Qwen architecture.
