Application (pre-grant publication)
Generation of Audio Including Emotionally Expressive Synthesized Content
- Number
- 20210090549
- Published
- 2021-03-25
- Filed
- 2019-09-23
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Lombardo; Salvator D., Kumar; Komath Naveen, Fidaleo; Douglas A.
- CPC
- G10L25/18; G10L13/047; G10L13/08; G10L13/033
- Verdict
- Low Notable software
- Source
- Google Patents · FreePatentsOnline
The keeper's note
Emotionally expressive speech synthesis for interactive characters.
Abstract
An audio processing system for generating audio including emotionally expressive synthesized content includes a computing platform having a hardware processor and a memory storing a software code including a trained neural network. The hardware processor is configured to execute the software code to receive an audio sequence template including one or more audio segment(s) and an audio gap, and to receive data describing one or more words for insertion into the audio gap. The hardware processor is configured to further execute the software code to use the trained neural network to generate an integrated audio sequence using the audio sequence template and the data, the integrated audio sequence including the one or more audio segment(s) and at least one synthesized word corresponding to the one or more words described by the data.
Background
BACKGROUND
The development of machine learning models for speech synthesis of emotionally expressive voices is challenging due to extensive variability in speaking styles. For example, the same word can be enunciated within a sentence in a variety of different ways to elicit unique characteristics, such as the emotional state of the speaker. As a result, training a successful model to generate a full sentence of speech typically requires a very large dataset, such as twenty hours or more of prerecorded speech.
Even when conventional neural speech generation models are successful, the speech they generate is often not emotionally expressive due at least in part to the fact that the training objective employed in conventional solutions is regression to the mean. Such a regression to the mean training objective encourages the conventional model to output a “most likely” averaged utterance, which tends not to sound convincing to the human ear. Consequently, expressive speech synthesis is usually not successful and remains a largely unsolved problem in the art. SUMMARY
There are provided systems and methods for generating audio including emotionally expressive synthesized content, substantially as shown in and/or described in connection with at least one of the figures, and as set forth more completely in the claims.