First-hand information for everyone
A team of researchers introduces SenseNova-U1.5, an 8‑billion‑parameter multimodal model that is built to understand, reason about, and gene
This paper introduces VBVR-Pro, a testbed that trains and measures “native visual reasoning.” That means the system uses images and videos t
Researchers introduce MMDiff, a new way to find and change the internal visual features that drive multimodal large language models (MLLMs).
This paper introduces PhiZero, a new way to model how the physical world unfolds in videos. Instead of directly predicting pixels for future
This paper introduces ClinFusion, a multimodal large language model (MLLM) designed to understand medical images the way clinicians do. The
This paper introduces VLM-IE3D, a way to give vision-language models a stronger sense of 3D space using only ordinary RGB video. The authors
Researchers introduce GigaPath-Flash and GigaTIME-Flash, two compact foundation models for pathology that aim to make whole-slide image anal
This paper introduces VideoRAE, a method that converts the internal features of large, frozen video encoders into compact representations th
This paper tackles a hard problem in digital breast tomosynthesis (DBT), an imaging method that makes a 3D picture of the breast from a smal
Medicine uses many types of information at once, such as images and text. Researchers from Yale and collaborators built MedPMC, an automated