Outer Rim Archives
Archives · 2018 · 9911218

Granted patent

Systems and methods for speech animation using visemes with phonetic boundary context

Number
9911218
Published
2018-03-06
Filed
2015-12-01
Assignee
Disney Enterprises, Inc.
Inventors
Theobald; Barry-John; Meyerhofer; Margaret; Matthews; Iain; Taylor; Sarah
CPC
G06T13/80; G06T13/205; G06T13/40; G10L21/10; G10L15/187
Verdict
Medium Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

Viseme-based speech animation.

Abstract

Speech animation may be performed using visemes with phonetic boundary context. A viseme unit may comprise an animation that simulates lip movement of an animated entity. Individual ones of the viseme units may correspond to one or more complete phonemes and phoneme context of the one or more complete phonemes. Phoneme context may include a phoneme that is adjacent to the one or more complete phonemes that correspond to a given viseme unit. Potential sets of viseme units that correspond with individual phoneme string portions may be determined. One of the potential sets of viseme units may be selected for individual ones of the phoneme string portions based on a fit metric that conveys a match between individual ones of the potential sets and the corresponding phoneme string portion.

Background

FIELD OF THE DISCLOSURE(1) This disclosure relates to speech animation using visemes with phonetic boundary context.BACKGROUND(2) Speech animation may require moving a jaw, lips, teeth and/or tongue of a facial model in synchrony with some accompanying audio, sometimes referred to as lip-syncing. Some approaches to speech animation may use visual movement parameters of the jaw, lips, teeth, tongue, and/or other facial features that represent speech sounds. Some techniques may use machine learning or probabilistic modeling techniques, such as hidden Markov models (HMMs) and/or hidden semi-Markov models (HSMMs). The models may be based on phonemes, which describe the acoustic sounds of a language.SUMMARY(3) One aspect of the disclosure relates to a system configured for speech animation using visemes with phonetic boundary context. Phonetic boundary context may account for viseme unit boundaries that partially span a phoneme. The introduction of phonetic context may improve the way in which viseme units may be selected for an input phoneme strings in real time or near real time. The improvements may include processing load reduction, combinations of viseme units producing facial movement which is smoother and more closely resembles human facial movement during speech, and/or other improvements. Individual viseme units may be usable for one or more animation entities (e.g., animated characters) using an underlying mesh or rig that defines facial feature movement of the entity. T

Claims

1. A system for speech animation using visemes with phonetic boundary context, the system comprising: one or more physical processors configured by machine-readable instructions to: obtain phoneme strings, the obtained phoneme strings including a first phoneme string, the first phoneme string including a first phoneme string portion; access information defining viseme units, the viseme units comprising animations that simulate lip movement of an animated entity, the viseme units including a first set of viseme units that simulate lip movement for one or more complete phonemes and a second set of viseme units that simulate lip movement for one or more partial phonemes, an individual partial phoneme spanning one of a beginning portion, a middle portion, or an end portion of an individual complete phoneme without inclusion of a remaining part of the individual complete phoneme; determine potential sets of viseme units that correspond with the first phoneme string portion, wherein different ones of the potential sets of viseme units form different viseme strings that define different animations of lip movement to simulate the first phoneme string portion, the potential sets of viseme units including a first potential set and a second potential set; and select one of the potential sets of viseme units based on a fit metric that conveys a match between individual ones of the potential sets and the first phoneme string portion. 14. A method speech animation using visemes with phonetic boundary context, the method being implemented in a computer system including one or more physical processors and storage media storing machine-readable instructions, the method comprising: obtaining phoneme strings, including obtaining a first phoneme string, the first phoneme string including a first phoneme string portion; accessing information defining viseme units, the viseme units comprising animations that simulate lip movement of an animated entity, the viseme units including a first set of viseme units that simulate lip movement for one or more complete phonemes and a second set of viseme units that simulate lip movement for one or more partial phonemes, an individual partial phoneme spanning one of a beginning portion, a middle portion, or an end portion of an individual complete phoneme without inclusion of a remaining part of the individual complete phoneme; determining potential sets of viseme units that correspond with the first phoneme string portion, wherein different ones of the potential sets of viseme units form different viseme strings that define different animations of lip movement to simulate the first phoneme string portion, including determining a first potential set and a second potential set; and selecting one of the potential sets of viseme units based on a fit metric that conveys a match between individual ones of the potential sets and the first phoneme string portion.