arXiv News

First-hand information for everyone

Site

AboutPrivacy PolicyTerms of Use

© 2026 arXiv News

arXiv News

EnglishJapanese

Switch language

EnglishJapanese
Loading account…

First-hand information for everyone

Latest
PhiZero teaches a model a compact "physical language" to predict how scenes changeClinFusion: a vision-focused multimodal LLM that reads 2D and 3D medical imagesVLM-IE3D teaches vision-language models 3D understanding from RGB video using implicit and explicit geometrySmaller, faster pathology AI models that keep most of the original performanceVideoRAE: turning frozen video encoders into compact latents for faster, better video generationExact data enforcement and calibrated uncertainty for limited-angle breast tomosynthesisMedPMC turns 6.1 million PubMed Central articles into 11 million medical image–text pairs for training multimodal modelsResearchers reframe computer vision as a single system that generates text and images from instructionsBenchX: 85,355 CT scans show tumor‑detection AI can fail for underrepresented groups and imaging protocolsPHAST-Net: a neural method that produces cleaner, high-resolution time–frequency views of soundPhiZero teaches a model a compact "physical language" to predict how scenes changeClinFusion: a vision-focused multimodal LLM that reads 2D and 3D medical imagesVLM-IE3D teaches vision-language models 3D understanding from RGB video using implicit and explicit geometrySmaller, faster pathology AI models that keep most of the original performanceVideoRAE: turning frozen video encoders into compact latents for faster, better video generationExact data enforcement and calibrated uncertainty for limited-angle breast tomosynthesisMedPMC turns 6.1 million PubMed Central articles into 11 million medical image–text pairs for training multimodal modelsResearchers reframe computer vision as a single system that generates text and images from instructionsBenchX: 85,355 CT scans show tumor‑detection AI can fail for underrepresented groups and imaging protocolsPHAST-Net: a neural method that produces cleaner, high-resolution time–frequency views of sound

Today's Briefing

Friday, July 31, 2026
AllArtificial IntelligenceMachine LearningNatural Language ProcessingComputer VisionRoboticsCryptographyPhysicsMathematics
Computer VisionFeatured briefing

PhiZero teaches a model a compact "physical language" to predict how scenes change

This paper introduces PhiZero, a new way to model how the physical world unfolds in videos. Instead of directly predicting pixels for future

July 31, 2026EN2 min read
Read full article

Latest Research

Artificial Intelligence
July 28, 2026

ClinFusion: a vision-focused multimodal LLM that reads 2D and 3D medical images

This paper introduces ClinFusion, a multimodal large language model (MLLM) designed to understand medical images the way clinicians do. The

EN
2 min read
Artificial Intelligence
July 24, 2026

VLM-IE3D teaches vision-language models 3D understanding from RGB video using implicit and explicit geometry

This paper introduces VLM-IE3D, a way to give vision-language models a stronger sense of 3D space using only ordinary RGB video. The authors

EN
2 min read
Artificial Intelligence
July 21, 2026

Smaller, faster pathology AI models that keep most of the original performance

Researchers introduce GigaPath-Flash and GigaTIME-Flash, two compact foundation models for pathology that aim to make whole-slide image anal

EN
2 min read
Advertisement
Computer Vision
July 16, 2026

VideoRAE: turning frozen video encoders into compact latents for faster, better video generation

This paper introduces VideoRAE, a method that converts the internal features of large, frozen video encoders into compact representations th

EN
2 min read
Computer Vision
July 15, 2026

Exact data enforcement and calibrated uncertainty for limited-angle breast tomosynthesis

This paper tackles a hard problem in digital breast tomosynthesis (DBT), an imaging method that makes a 3D picture of the breast from a smal

EN
2 min read
Machine Learning
July 9, 2026

MedPMC turns 6.1 million PubMed Central articles into 11 million medical image–text pairs for training multimodal models

Medicine uses many types of information at once, such as images and text. Researchers from Yale and collaborators built MedPMC, an automated

EN
2 min read
Computer Vision
July 8, 2026

Researchers reframe computer vision as a single system that generates text and images from instructions

This paper treats computer vision as a single “multimodal generation” problem. Instead of building a different model for each vision task, t

EN
2 min read
Computer Vision
June 24, 2026

BenchX: 85,355 CT scans show tumor‑detection AI can fail for underrepresented groups and imaging protocols

This paper introduces BenchX, a large, open benchmark built to test how well artificial intelligence (AI) models detect and locate tumors in

EN
2 min read
Computer Vision
June 23, 2026

PHAST-Net: a neural method that produces cleaner, high-resolution time–frequency views of sound

Researchers introduce PHAST-Net, a neural network that turns a set of wavelet-based measurements into clean time–frequency pictures of sound

EN
2 min read
Next page of briefings