Outer Rim Archives
Archives · 2021 · 11062692

Granted patent

Generation of audio including emotionally expressive synthesized content

Number
11062692
Published
2021-07-13
Filed
2019-09-23
Assignee
Disney Enterprises, Inc.
Inventors
Lombardo; Salvator D., Kumar; Komath Naveen, Fidaleo; Douglas A.
CPC
G10L25/18; G10L13/033; G10L13/047; G10L13/08
Verdict
Low Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

Emotionally expressive speech synthesis for interactive characters (granted).

Abstract

An audio processing system for generating audio including emotionally expressive synthesized content includes a computing platform having a hardware processor and a memory storing a software code including a trained neural network. The hardware processor is configured to execute the software code to receive an audio sequence template including one or more audio segment(s) and an audio gap, and to receive data describing one or more words for insertion into the audio gap. The hardware processor is configured to further execute the software code to use the trained neural network to generate an integrated audio sequence using the audio sequence template and the data, the integrated audio sequence including the one or more audio segment(s) and at least one synthesized word corresponding to the one or more words described by the data.

Background

BACKGROUND (1) The development of machine learning models for speech synthesis of emotionally expressive voices is challenging due to extensive variability in speaking styles. For example, the same word can be enunciated within a sentence in a variety of different ways to elicit unique characteristics, such as the emotional state of the speaker. As a result, training a successful model to generate a full sentence of speech typically requires a very large dataset, such as twenty hours or more of prerecorded speech. (2) Even when conventional neural speech generation models are successful, the speech they generate is often not emotionally expressive due at least in part to the fact that the training objective employed in conventional solutions is regression to the mean. Such a regression to the mean training objective encourages the conventional model to output a “most likely” averaged utterance, which tends not to sound convincing to the human ear. Consequently, expressive speech synthesis is usually not successful and remains a largely unsolved problem in the art. SUMMARY (3) There are provided systems and methods for generating audio including emotionally expressive synthesized content, substantially as shown in and/or described in connection with at least one of the figures, and as set forth more completely in the claims.

Claims

1. An audio processing system comprising: a computing platform including a hardware processor and a system memory; a software code stored in the system memory, the software code including a trained neural network; the hardware processor configured to execute the software code to: receive an audio sequence template including at least one audio segment and an audio gap; receive data describing at least one word for insertion into the audio gap; and use the trained neural network to generate an integrated audio sequence using the audio sequence template and the data, the integrated audio sequence including the at least one audio segment and at least one synthesized word corresponding to the at least one word described by the data. || 11. A method for use by an audio processing system including a computing platform having a hardware processor and a system memory storing a software code including a trained neural network, the method comprising: receiving, by the software code executed by the hardware processor, an audio sequence template including at least one audio segment and an audio gap; receiving, by the software code executed by the hardware processor, data describing at least one word for insertion into the audio gap; and using the trained neural network, by the software code executed by the hardware processor, to generate an integrated audio sequence using the audio sequence template and the data, the integrated audio sequence including the at least one audio segment and at least one synthesized word corresponding to the at least one word described by the data.