Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reasoning images, Visual CoT provides an intuitive mechanism for visual foresight but introduces substantial inference overhead, which is particularly challenging for real-time applications. Research into internalized visual thinking aims to address this overhead while maintaining the reasoning capabilities of Visual CoT for proactive video reasoning.
Read the original source at Apple Machine Learning Research ↗
