First-hand information for everyone
This paper introduces PhiZero, a new way to model how the physical world unfolds in videos. Instead of directly predicting pixels for future
This paper introduces ClinFusion, a multimodal large language model (MLLM) designed to understand medical images the way clinicians do. The
This paper introduces VLM-IE3D, a way to give vision-language models a stronger sense of 3D space using only ordinary RGB video. The authors
Researchers introduce GigaPath-Flash and GigaTIME-Flash, two compact foundation models for pathology that aim to make whole-slide image anal
This paper introduces VideoRAE, a method that converts the internal features of large, frozen video encoders into compact representations th
This paper tackles a hard problem in digital breast tomosynthesis (DBT), an imaging method that makes a 3D picture of the breast from a smal
Medicine uses many types of information at once, such as images and text. Researchers from Yale and collaborators built MedPMC, an automated
This paper treats computer vision as a single “multimodal generation” problem. Instead of building a different model for each vision task, t
This paper introduces BenchX, a large, open benchmark built to test how well artificial intelligence (AI) models detect and locate tumors in
Researchers introduce PHAST-Net, a neural network that turns a set of wavelet-based measurements into clean time–frequency pictures of sound