Outer Rim Archives
Archives · 2026 · 20260253579

Application (pre-grant publication)

MEASUREMENT OF PROSODY SIMILARITY USING A PROSODY ENCODER

Number
20260253579
Published
2026-08-27
Filed
2025-02-24
Assignee
Disney Enterprises, Inc.
Inventors
Kumar; Komath Naveen, Doggett; Erika, UM; SeYun
CPC
G10L13/10; G10L25/30; G10L25/51
Verdict
Set aside dropped in weekly review
In edition
2026-W36
Source
Google Patents · FreePatentsOnline

The keeper's note

a method inputs a first speech sample into a prosody encoder, wherein the prosody encoder is trained as part of a text-to-speech model.

Abstract

In some embodiments, a method inputs a first speech sample into a prosody encoder, wherein the prosody encoder is trained as part of a text-to-speech model. A second speech sample is input into the prosody encoder. The prosody encoder generates a first representation of the first speech sample and a second representation of the second speech sample. The method compares the first representation and the second representation to determine a metric value and determines an action to perform based on the metric value.

Background

BACKGROUND

The determination of the similarity of speech samples may require a large amount of resources. Subjective comparison may be used where an expert listener may listen to the speech samples to subjectively determine similarity. However, the comparison may not be performed at scale. Other methods may be used to attempt to determine similarity. For example, speaker identity models may be used, which specialize in identifying the speaker. However, this is a narrow use case that only identifies the speaker.

Claims

1. A method comprising: inputting a first speech sample into a prosody encoder, wherein the prosody encoder is trained as part of a text-to-speech model; inputting a second speech sample into the prosody encoder; generating, using the prosody encoder, a first representation of the first speech sample and a second representation of the second speech sample; comparing the first representation and the second representation to determine a metric value; and determining an action to perform based on the metric value. || 18. A non-transitory computer-readable storage medium having stored thereon computer executable instructions, which when executed by a computing device, cause the computing device to be operable for: inputting a first speech sample into a prosody encoder, wherein the prosody encoder is trained as part of a text-to-speech model; inputting a second speech sample into the prosody encoder; generating, using the prosody encoder, a first representation of the first speech sample and a second representation of the second speech sample; comparing the first representation and the second representation to determine a metric value; and determining an action to perform based on the metric value. || 20. An apparatus comprising: one or more computer processors; and a computer-readable storage medium comprising instructions for controlling the one or more computer processors to be operable for: inputting a first speech sample into a prosody encoder, wherein the prosody encoder is trained as part of a text-to-speech model; inputting a second speech sample into the prosody encoder; generating, using the prosody encoder, a first representation of the first speech sample and a second representation of the second speech sample; comparing the first representation and the second representation to determine a metric value; and determining an action to perform based on the metric value.