First-hand information for everyone
Researchers introduce MP-Bench, the first benchmark made to test voice agents as active participants in multiparty conversations. The benchm
This paper argues that small Transformer models fail on some compositional generalisation tests not because the models are fundamentally bad
Researchers tested open-source Whisper speech tools on real-world YouTube videos in seven languages and found that automatic transcripts are
Biologists often cannot test every possible genetic perturbation in the lab. CRISPR screens ask which genes or edits change a cell’s behavio
Speech large language models (SpeechLLMs) work directly from audio. They are faster than systems that first convert speech to text and they
The paper asks a simple but important question: when a language model gives a probability for an event, do its probabilities hang together i
This paper describes a focused post‑training process that made language models much stronger at competitive programming. Competitive program
This paper introduces OmniScientist, an AI system that tries to do end-to-end scientific work starting from raw evidence. Instead of only re
Researchers introduce MMDiff, a new way to find and change the internal visual features that drive multimodal large language models (MLLMs).
Researchers compared two ways large language models (LLMs) call external tools. The first way is the common “JSON tool calling” where the mo