Outer Rim Archives
Archives · 2025 · 20250372116

Application (pre-grant publication)

Cross-Language Voice Similarity Analysis

Number
20250372116
Published
2025-12-04
Filed
2024-08-28
Assignee
Disney Enterprises, Inc.
Inventors
Beard; Audrey Coyote Aura et al.
CPC
G10L25/51; G10L15/22; G10L17/22; G10L25/30
Verdict
Set aside localization/dubbing voice analysis, business
Source
Google Patents · FreePatentsOnline

Abstract

A system includes a hardware processor and a memory storing a cross-language voice similarity analyzer (analyzer). The hardware processor executes the analyzer to generate an embedding vector representation of an audio sample of a human voice in a feature space including existing embedding vectors corresponding respectively to different reference voices, decompose the embedding vector representation to identify a linear or non-linear combination of vocal component vectors corresponding to the human voice, each vocal component vector representing a respective predetermined voice characteristic descriptor, increase the dimensionality of the linear or non-linear combination of the vocal component vectors to match the dimensionality of the embedding vector representation to provide a reconstructed embedding vector representation of the human voice, and identify, by comparing the reconstructed embedding vector representation with one or more of the existing embedding vectors, one of the reference voices as a match for the audio sample of the human voice.

Background

BACKGROUND

Major film and television (TV) studios typically produce large amounts of audiovisual (AV) content (e.g., feature films, episodic TV content, and the like) in a single language for their primary home market. To maximize the value of this content and compete internationally, these studios often undertake a meticulous localization process, whereby a given piece of AV content is modified to be more relevant and comprehensible to consumers in a target foreign market.

One of the most common forms of localization is dubbing, in which all of the source language dialog and vocal musical performances are replaced with appropriate dialog and vocals in the target foreign language. The voice talent cast for these localized versions are often chosen because their voice or affectation matches closely with that of the original version. Historically, this voice casting process has been largely manual, iterative, expensive, and biased towards previously cast talent.

Voice casting for regionally localized content, using conventional methods, is a highly manual process that requires personnel within international casting departments to search through many auditions to find similar sounding voices. This is a process that has traditionally been managed by a few human experts who hold years of embodied knowledge, resulting in the risk of bias based on established relationships with or preferences for certain talent but not others, and brittleness based on the chance

Claims

1. A system comprising: a computing platform including a hardware processor; and a memory storing a cross-language voice similarity analyzer; the hardware processor configured to execute the cross-language voice similarity analyzer to: generate, using an audio sample of a human voice, an embedding vector representation of the human voice in a multi-dimensional feature space including a plurality of existing embedding vectors each corresponding respectively to a reference voice of a plurality of reference voices; decompose the embedding vector representation of the human voice to identify a linear combination of vocal component vectors corresponding to the human voice or a non-linear combination of vocal component vectors corresponding to the human voice, each of the vocal component vectors representing a respective one of a plurality of predetermined voice characteristic descriptors; increase a dimensionality of the linear combination of the vocal component vectors or the non-linear combination of the vocal component vectors to match a dimensionality of the embedding vector representation of the human voice to provide a reconstructed embedding vector representation of the human voice; and identify, by comparing the reconstructed embedding vector representation of the human voice with one or more of the plurality of existing embedding vectors, one of the plurality of reference voices as a match for the audio sample of the human voice. || 11. A method for use by a system including a hardware processor and a memory storing a cross-language voice similarity analyzer, the method comprising: generating, by the cross-language voice similarity analyzer executed by the hardware processor and using an audio sample of a human voice, an embedding vector representation of the human voice in a multi-dimensional feature space including a plurality of existing embedding vectors each corresponding respectively to a reference voice of a plurality of reference voices; decomposing, by the cross-language voice similarity analyzer executed by the hardware processor, the embedding vector representation of the human voice to identify a linear combination of vocal component vectors corresponding to the human voice or a non-linear combination of vocal component vectors corresponding to the human voice, each of the vocal component vectors representing a respective one of a plurality of predetermined voice characteristic descriptors; increasing a dimensionality of the linear combination of the vocal component vectors or the non-linear combination of the vocal component vectors to match a dimensionality of the embedding vector representation of the human voice to provide a reconstructed embedding vector representation of the human voice; and identifying, by the cross-language voice similarity analyzer executed by the hardware processor through comparison of the reconstructed embedding vector representation of the human voice with one or more of the plurality of existing embedding vectors, one of the plurality of reference voices as a match for the audio sample of the human voice.