Microsoft's VALL-E AI can mimic any voice from a short audio sample

TL;DR

It's derived from Meta's AI-powered compression neural net Encodec, generating audio from text input and short samples from the target speaker.In a paper, researchers describe how they trained VALL-E on 60,000 hours of English language speech from 7,000-plus speakers on Meta's LibriLight audio library.If that's the case, it uses the training data to infer what the target speaker would sound like if speaking the desired text input.For each phrase they want the AI to "speak," they have a three-second prompt from the speaker to imitate, a "ground truth" of the same speaker saying another phrase for comparison, a "baseline" conventional text-to-speech synthesis and the VALL-E sample at the end."Since VALL-E could synthesize speech that maintains speaker identity, it may carry potential risks in misuse of the model, such as spoofing voice identification or impersonating," the company wrote in the "Broader impacts" section of its conclusion."

Like summarized versions? Support us on Patreon!