VLM-IE3D teaches vision-language models 3D understanding from RGB video using implicit and explicit geometry | arXiv News