arXiv News

First-hand information for everyone

Site

AboutPrivacy PolicyTerms of Use

© 2026 arXiv News

arXiv News

EnglishJapanese

Switch language

EnglishJapanese
Loading account…

First-hand information for everyone

Latest
MP-Bench tests how well voice assistants join group conversations — and finds they struggleMore kinds of grammatical building blocks help Transformers learn new sentence structuresWhisper-based transcription of YouTube videos: about 30% error across seven languages, reduced to ~20% with some fine-tuningLearning from past CRISPR screens to pick better experiments: a new benchmark and methodRetroThinker lets a speech-based language model revise its own step-by-step reasoning while you speakResearchers test whether language models give internally consistent probability forecasts using a "Dutch book" checkSpecialized post‑training lets an AI beat the top human score at the IOI coding contestOmniScientist: an AI that reads raw scientific data and writes full papersMMDiff: a method to find and control visual features inside multimodal large language modelsWriting code beats rigid JSON calls for many modern LLMs, study findsMP-Bench tests how well voice assistants join group conversations — and finds they struggleMore kinds of grammatical building blocks help Transformers learn new sentence structuresWhisper-based transcription of YouTube videos: about 30% error across seven languages, reduced to ~20% with some fine-tuningLearning from past CRISPR screens to pick better experiments: a new benchmark and methodRetroThinker lets a speech-based language model revise its own step-by-step reasoning while you speakResearchers test whether language models give internally consistent probability forecasts using a "Dutch book" checkSpecialized post‑training lets an AI beat the top human score at the IOI coding contestOmniScientist: an AI that reads raw scientific data and writes full papersMMDiff: a method to find and control visual features inside multimodal large language modelsWriting code beats rigid JSON calls for many modern LLMs, study finds

Today's Briefing

Monday, September 14, 2026
AllArtificial IntelligenceMachine LearningNatural Language ProcessingComputer VisionRoboticsCryptographyPhysicsMathematics
Artificial IntelligenceFeatured briefing

MP-Bench tests how well voice assistants join group conversations — and finds they struggle

Researchers introduce MP-Bench, the first benchmark made to test voice agents as active participants in multiparty conversations. The benchm

September 14, 2026EN2 min read
Read full article

Latest Research

Natural Language Processing
September 14, 2026

More kinds of grammatical building blocks help Transformers learn new sentence structures

This paper argues that small Transformer models fail on some compositional generalisation tests not because the models are fundamentally bad

EN
2 min read
Natural Language Processing
September 12, 2026

Whisper-based transcription of YouTube videos: about 30% error across seven languages, reduced to ~20% with some fine-tuning

Researchers tested open-source Whisper speech tools on real-world YouTube videos in seven languages and found that automatic transcripts are

EN
2 min read
Artificial Intelligence
September 11, 2026

Learning from past CRISPR screens to pick better experiments: a new benchmark and method

Biologists often cannot test every possible genetic perturbation in the lab. CRISPR screens ask which genes or edits change a cell’s behavio

EN
2 min read
Advertisement
Artificial Intelligence
September 11, 2026

RetroThinker lets a speech-based language model revise its own step-by-step reasoning while you speak

Speech large language models (SpeechLLMs) work directly from audio. They are faster than systems that first convert speech to text and they

EN
2 min read
Artificial Intelligence
September 3, 2026

Researchers test whether language models give internally consistent probability forecasts using a "Dutch book" check

The paper asks a simple but important question: when a language model gives a probability for an event, do its probabilities hang together i

EN
2 min read
Artificial Intelligence
September 3, 2026

Specialized post‑training lets an AI beat the top human score at the IOI coding contest

This paper describes a focused post‑training process that made language models much stronger at competitive programming. Competitive program

EN
2 min read
Artificial Intelligence
August 14, 2026

OmniScientist: an AI that reads raw scientific data and writes full papers

This paper introduces OmniScientist, an AI system that tries to do end-to-end scientific work starting from raw evidence. Instead of only re

EN
2 min read
Artificial Intelligence
August 11, 2026

MMDiff: a method to find and control visual features inside multimodal large language models

Researchers introduce MMDiff, a new way to find and change the internal visual features that drive multimodal large language models (MLLMs).

EN
2 min read
Natural Language Processing
August 7, 2026

Writing code beats rigid JSON calls for many modern LLMs, study finds

Researchers compared two ways large language models (LLMs) call external tools. The first way is the common “JSON tool calling” where the mo

EN
2 min read
Next page of briefings