Outer Rim Archives
Archives · 2024 · 11983808

Granted patent

Conversation-driven character animation

Number
11983808
Published
2024-05-14
Filed
2022-08-24
Assignee
Disney Enterprises, Inc.
Inventors
Tiwari; Sanchita et al.
CPC
H04L51/02; G06F40/35; G06N20/00; G06T13/80; G06T13/40
Verdict
Low Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

Conversation-driven interactive character animation technique (granted).

Abstract

A system for producing conversation-driven character animation includes a computing platform having processing hardware and a system memory storing software code, the software code including multiple trained machine learning (ML) models. The processing hardware executes the software code to obtain a conversation understanding feature set describing a present state of a conversation between a digital character and a system user, and to generate an inference, using at least a first trained ML model of the multiple trained ML models and the conversation understanding feature set, the inference including labels describing a predicted next state of a scene within the conversation. The processing hardware further executes the software code to produce, using at least a second trained ML model of the multiple trained ML models and the labels, an animation stream of the digital character participating in the predicted next state of the scene within the conversation.

Background

BACKGROUND (1) Advances in artificial intelligence (AI) have led to the generation of a variety of digital characters, such as avatars for example, that simulate social interaction. However, conventionally generated AI digital characters typically project a single synthesized persona that tends to lack character and naturalness. (2) In contrast to conventional interaction simulations by AI digital characters, natural interactions between human beings are more nuanced, varied, and dynamic. For example, interactions between humans are typically responsive to a variety of factors including environmental features such as location, weather, and lighting, the respective goals of the interaction partners, their respective emotional states, and the content of a conversation between them. That is to say, typical shortcomings of AI digital character interactions include their failure to integrate verbal communications with non-verbal cues arising from, complementing, and enhancing the content of those verbal communications. Thus, there is a need in the art for an animation solution capable of producing an animation sequence of a digital character participating in an interaction that is dynamically conversation-driven in real-time by the dialogue between the digital character and an interaction partner such as a human user.

Claims

1. A system comprising: a computing platform having a processing hardware, a plurality of sensors and a system memory storing a software code, the software code including a plurality of trained machine learning (ML) models; the processing hardware configured to execute the software code to: generate a conversation understanding feature set using at least a first trained ML model of the plurality of trained ML models and sensor data received from the plurality of sensors, the conversation understanding feature set describing a present state of a conversation between a digital character and a system user; generate an inference, using at least a second trained ML model of the plurality of trained ML models and the conversation understanding feature set, the inference comprising a plurality of labels describing a predicted next state of a scene within the conversation; and produce, using at least a third trained ML model of the plurality of trained ML models and the plurality of labels, an animation stream of the digital character participating in the predicted next state of the scene within the conversation. || 11. A method for use by a system including a computing platform having a processing hardware, a plurality of sensors and a system memory storing a software code, the software code including a plurality of trained machine learning (ML) models: generating, by the software code executed by the processing hardware, a conversation understanding feature set using at least a first trained ML model of the plurality of trained ML models and sensor data received from the plurality of sensors, the conversation understanding feature set describing a present state of a conversation between a digital character and a system user; generating an inference, by the software code executed by the processing hardware and using at least a second trained ML model of the plurality of trained ML models and the conversation understanding feature set, the inference comprising a plurality of labels describing a predicted next state of a scene within the conversation; and producing, by the software code executed by the processing hardware and using at least a third trained ML model of the plurality of trained ML models and the plurality of labels, an animation stream of the digital character participating in the predicted next state of the scene within the conversation. || 22. A method for use by a system including a computing platform having a processing hardware and a system memory storing a software code, the software code including a plurality of trained machine learning (ML) models: obtaining, by the software code executed by the processing hardware, a conversation understanding feature set describing a present state of a conversation between a digital character and a system user; generating an inference, by the software code executed by the processing hardware and using at least a first trained ML model of the plurality of trained ML models and the conversation understanding feature set, the inference comprising a plurality of labels transformed into two-dimensional (2D) or three-dimensional (3D) skeletal keypoints for a skeleton model of the digital character, the plurality of labels describing a predicted next state of a scene within the conversation; and producing, by the software code executed by the processing hardware and using at least a second trained ML model of the plurality of trained ML models and the plurality of labels, an animation stream of the digital character participating in the predicted next state of the scene within the conversation.