RePlay: a speech system that retrieves and plays pre-recorded lines to cut response delay
Many voice systems must not only say the right words, but say them exactly and in a fixed voice. Generative speech models are fast but can’t guarantee an exact prerecorded performance. Traditional cascaded systems (which use automatic speech recognition, or ASR, then a text model and text-to-speech) can enforce exact lines but are slower. This paper presents RePlay, a hybrid approach that retrieves a matching prerecorded line and plays it as soon as the system turn starts, aiming for the speed of end-to-end models with the control of scripted systems.
The researchers built RePlay by adapting an existing spoken model called PersonaPlex. They probed the model’s internal state to find when and where the upcoming response is already encoded. They found that after a special turn-start token (called <EPAD>) is fed back into the model, the response information becomes recoverable around layer 15 of the 32-layer language model. Based on this, RePlay keeps only the first 16 layers, and adds two small heads at layer 15: one to detect the turn start and one to map the hidden state to an embedding that can be used to pick a prerecorded line.
At run time, RePlay detects the start of the system turn and then scores all candidate lines by cosine similarity to the projected hidden state. The highest-scoring line’s pre-recorded audio is played immediately. Because the system uses frozen text embeddings for the candidate lines, new lines can be added without retraining. The design removes the need for ASR or text-to-speech during the response, and it halves the model compute compared with the full 32-layer model.
The team tested RePlay on synthetic multi-turn dialogues and on a job-interview scenario with a scripted character named Nobu. In simulated interviews, RePlay reached a median latency of 383 milliseconds, which the authors report is three to seven times faster than cascaded ASR+LLM systems that achieved similar dialogue quality. In a user study, participants preferred RePlay in 63% of ratings versus 12% for a fast cascade using a small large language model (p = 0.008). Against a slower cascade with a stronger LLM, preferences were 46% versus 21%, a difference the authors call not statistically significant. The authors also note RePlay has lower exact-line accuracy than cascaded systems.