arXiv News

First-hand information for everyone

Site

AboutPrivacy PolicyTerms of Use

© 2026 arXiv News

arXiv News

EnglishJapanese

Switch language

EnglishJapanese
Loading account…

First-hand information for everyone

Latest
SenseNova-U1.5: an 8-billion-parameter model that reads and creates images end-to-endVBVR-Pro: a testbed that teaches machines to reason by creating images and videosMMDiff: a method to find and control visual features inside multimodal large language modelsPhiZero teaches a model a compact "physical language" to predict how scenes changeClinFusion: a vision-focused multimodal LLM that reads 2D and 3D medical imagesVLM-IE3D teaches vision-language models 3D understanding from RGB video using implicit and explicit geometrySmaller, faster pathology AI models that keep most of the original performanceVideoRAE: turning frozen video encoders into compact latents for faster, better video generationExact data enforcement and calibrated uncertainty for limited-angle breast tomosynthesisMedPMC turns 6.1 million PubMed Central articles into 11 million medical image–text pairs for training multimodal modelsSenseNova-U1.5: an 8-billion-parameter model that reads and creates images end-to-endVBVR-Pro: a testbed that teaches machines to reason by creating images and videosMMDiff: a method to find and control visual features inside multimodal large language modelsPhiZero teaches a model a compact "physical language" to predict how scenes changeClinFusion: a vision-focused multimodal LLM that reads 2D and 3D medical imagesVLM-IE3D teaches vision-language models 3D understanding from RGB video using implicit and explicit geometrySmaller, faster pathology AI models that keep most of the original performanceVideoRAE: turning frozen video encoders into compact latents for faster, better video generationExact data enforcement and calibrated uncertainty for limited-angle breast tomosynthesisMedPMC turns 6.1 million PubMed Central articles into 11 million medical image–text pairs for training multimodal models

Today's Briefing

Monday, September 14, 2026
AllArtificial IntelligenceMachine LearningNatural Language ProcessingComputer VisionRoboticsCryptographyPhysicsMathematics
Computer VisionFeatured briefing

SenseNova-U1.5: an 8-billion-parameter model that reads and creates images end-to-end

A team of researchers introduces SenseNova-U1.5, an 8‑billion‑parameter multimodal model that is built to understand, reason about, and gene

September 11, 2026EN2 min read
Read full article

Latest Research

Artificial Intelligence
August 27, 2026

VBVR-Pro: a testbed that teaches machines to reason by creating images and videos

This paper introduces VBVR-Pro, a testbed that trains and measures “native visual reasoning.” That means the system uses images and videos t

EN
2 min read
Artificial Intelligence
August 11, 2026

MMDiff: a method to find and control visual features inside multimodal large language models

Researchers introduce MMDiff, a new way to find and change the internal visual features that drive multimodal large language models (MLLMs).

EN
2 min read
Computer Vision
July 31, 2026

PhiZero teaches a model a compact "physical language" to predict how scenes change

This paper introduces PhiZero, a new way to model how the physical world unfolds in videos. Instead of directly predicting pixels for future

EN
2 min read
Advertisement
Artificial Intelligence
July 28, 2026

ClinFusion: a vision-focused multimodal LLM that reads 2D and 3D medical images

This paper introduces ClinFusion, a multimodal large language model (MLLM) designed to understand medical images the way clinicians do. The

EN
2 min read
Artificial Intelligence
July 24, 2026

VLM-IE3D teaches vision-language models 3D understanding from RGB video using implicit and explicit geometry

This paper introduces VLM-IE3D, a way to give vision-language models a stronger sense of 3D space using only ordinary RGB video. The authors

EN
2 min read
Artificial Intelligence
July 21, 2026

Smaller, faster pathology AI models that keep most of the original performance

Researchers introduce GigaPath-Flash and GigaTIME-Flash, two compact foundation models for pathology that aim to make whole-slide image anal

EN
2 min read
Computer Vision
July 16, 2026

VideoRAE: turning frozen video encoders into compact latents for faster, better video generation

This paper introduces VideoRAE, a method that converts the internal features of large, frozen video encoders into compact representations th

EN
2 min read
Computer Vision
July 15, 2026

Exact data enforcement and calibrated uncertainty for limited-angle breast tomosynthesis

This paper tackles a hard problem in digital breast tomosynthesis (DBT), an imaging method that makes a 3D picture of the breast from a smal

EN
2 min read
Machine Learning
July 9, 2026

MedPMC turns 6.1 million PubMed Central articles into 11 million medical image–text pairs for training multimodal models

Medicine uses many types of information at once, such as images and text. Researchers from Yale and collaborators built MedPMC, an automated

EN
2 min read
Next page of briefings