Fixing detection mistakes in sound-scene separation by checking alternative labels at inference time
This paper studies how to reduce errors that happen when systems first detect sounds and then try to extract them from a recording. In many pipelines, a mistake in the detection step—saying the wrong sound class is present—breaks the second step that isolates that sound. The authors build a multichannel detection-and-extraction system and test methods that, at inference time, propose and score alternative class labels and the corresponding separated sounds to recover from detection errors.
The system is a two-stage pipeline. A multi-label audio classifier (M2D) predicts which of 18 target sound classes are active in a four-channel mixture. Those predicted class labels are used as prompts for a target source extractor based on a multichannel variant of the TUSS model (Mc-TUSS). The task and data come from the DCASE 2025 Task 4 challenge, where each example has up to three simultaneous target events plus background noise and interference. The main evaluation metric is class-aware signal-to-distortion ratio improvement (CA-SDRi), which only rewards good separation when the predicted class label matches the true class.
To correct detection mistakes without retraining, the authors generate alternative label sets and corresponding source estimates at inference time. They compare two relabeling strategies that use the classifier’s confusion matrix: a “naive” strategy that replaces each predicted label by its top confusions (yielding up to 15 alternatives when three targets are present), and a “conditional” strategy that redistributes the classifier’s probabilities according to the confusion patterns and then scores all valid class combinations (they enumerate 987 combinations and keep the top candidates). Each candidate label set is scored by a fitness function and by models that judge the extracted sources. The paper tests two fitness functions—binary cross-entropy and mixture consistency—and three ways to judge source estimates: a multi-label classifier, a fine-tuned single-source classifier, and an off-the-shelf “audio judge.” They also validate an entropy-based gate that decides when to apply inference-time updates on a held-out set.