This AI Can Preserve A Person’s Voice With Just A Few Hours Of Recordings

Date:

VocaliD is a pioneering company dedicated to preserving and re-creating people’s voices using artificial intelligence. It has recently joined forces with Northeastern University in Boston to open a center for the service – called The Voice Preservation Clinic.

The researchers hope their efforts will change the lives of individuals who face losing the ability to speak due to illnesses, like throat cancer or motor neurone disease. They want to provide people with a means of maintaining their sense of identity. Being able to keep their voice even when it becomes impossible for them to self generate speech should help with that.

Before the collaboration, VocaliD provided the service and allowed people to record their voices from the comforts of their home. However, this was an inadequate method because most people lacked equipment for high-quality recordings, or they made recordings with background noise. Another problem was most people didn’t even know this option existed.

Prof. Rupal Patel, founder, and chief executive of VocaliD and the center’s lead researcher, said:

Oftentimes, they’ll come to us at the last minute. They don’t have enough time to bank their voice and they are also just so enveloped in their disease and then the surgery – that is very stressful.

AI can preserve someone's voice from a few recordings

That’s why the company teamed up with Northeastern, to make the technology more accessible to the masses and to provide the patients with a proper recording environment for good quality sound. They are calling this the “legacy project.”

How It Works

Step 1: The person “banks” their voice by recording themselves talking. The clinic provides the participant with poems, speeches, and short stories in a range of topics. The recordings take place in a special booth.

Patel said:

What we have them do is record about two to three hours of speech. From those recordings, we then are able to build an AI-generated voice engine, essentially, that sounds like them.

Step 2: The clinic takes the recordings and feeds them to machine-learning algorithms. The procedure is more sophisticated than taking a bunch of words, chopping them up, and then stringing them together. The AI-generated voice engine doesn’t just repeat the words that were recorded. It breaks down the sound of the voice and can say words – in the user’s voice – that were never recorded!

Step 3: The digital voice is then installed on the accompanying app that has been installed into the user’s phone or another device. Then, the person can type what they want to say, and the app will produce the audio of the sentences in the user’s voice.

Patel said:

With the cancer population, they have control of their hands, and they can communicate – but they want to communicate as themselves.

A Voice That Grows

The technology can even age a person’s voice so that it grows with them. Although, it isn’t possible yet to turn a child’s voice into an adult. There are also filters in development that will give a user more choices as far as how phrases are expressed.

Andrea D. Steffen
Andrea D. Steffen
I use the alphabet to paint words that become a beautiful and inspiring image in the reader's mind. I have a Bachelors in Architecture from FAU.

Share post:

Popular

The 12 Best Indoor Plants for Air Purification: Care, Cost, and Pet Safety

Indoor plants do far more than brighten a room....

DeepSeek Price Increase: New V4 Rates, Cache Economics, and DeepSeek Alternatives

Developers and enterprise teams worldwide were taken by surprise...

Transforming Properties with Professional Landscape and Event Lighting

Why do some properties command attention after dark while...

Why Extending a Song Is Harder Than Pressing Loop

A short piece of music can be exactly right...