Outer Rim Archives
Archives · 2024 · 20240386217

Application (pre-grant publication)

Entertainment Character Interaction Quality Evaluation and Improvement

Number
20240386217
Published
2024-11-21
Filed
2024-03-01
Assignee
Disney Enterprises, Inc.
Inventors
Doggett; Erika Varis et al.
CPC
G06F40/30; G06F40/40
Verdict
Low Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

Interaction-quality evaluation/improvement for interactive characters.

Abstract

A system includes a processor and a memory storing software code. The processor executes the software code to receive dialogue data identifying a character, a storyline including the character, and speech for the character intended to advance the storyline or achieve a goal, assess, using the dialogue data, quality assurance (QA) metrics of the speech including at least one of: (i) its fluency, (ii) its responsiveness to speech by an interaction partner of the character, (iii) its consistency with the goal, (iv) its consistency with a character profile of the character, or (v) or its consistency with a story-world of the storyline, and determine, using the QA metrics, whether the speech is suitable for advancing the storyline or achieving the goal. When determining determines that the speech is suitable, approve the speech. When determining determines that the speech is unsuitable, flag the speech as unsuitable.

Background

BACKGROUND

Recent work in generative language modeling has inspired explorations into character dialogue generation that is more flexible than conventional pre-authored dialogue tree approaches, and can generate a wide range of character responses quickly and easily. However, measuring the quality of those model generated responses is underdeveloped in the existing art, leaving designers to either use metrics unsuited for character interactions and missing key components of what makes the persona of a character distinctive, or alternatively to not use metrics at all and rely purely on the internal probabilistic values produced by the model with no external judgment.

Previous metrics for judging in-character consistency have sometimes used an entailment model, but fail to account for in-world consistency, i.e., whether the interactions of the character are consistent with the historical time and location of that character, or whether the character is staying consistent to a goal of an interaction. Moreover, existing work tends to be heavily focused on content-level metrics like toxicity and truthfulness. Outside of those content-level metrics, engineers and researchers have largely been limited to surface-level metrics such as grammar and semantics, and the current dominant metric used for evaluating large language models (i.e., perplexity), which corresponds to internal likelihood consistency for sentence structure, and is insufficient for judging any of the met

Claims

1. A system comprising: a hardware processor; and a memory storing a software code; the hardware processor configured to execute the software code to: receive dialogue data, the dialogue data identifying a character, a storyline including the character, and a speech for the character intended to at least one of advance the storyline or achieve a goal of the speech; assess, using the dialogue data, a plurality of quality assurance (QA) metrics of the speech, the plurality of QA metrics including at least one of: (i) a fluency of the speech, (ii) a responsiveness of the speech to speech by an interaction partner of the character, (iii) a consistency of the speech with the goal of the speech, (iv) a consistency of the speech with a character profile of the character, or (v) a consistency of the speech with a story-world of the storyline; determine, using the plurality of QA metrics, whether the speech is suitable for advancing the storyline; when determining determines that the speech is suitable for advancing the storyline or achieving the goal, approve the speech; and when determining determines that the speech is unsuitable for advancing the storyline or achieving the goal, flag the speech as being unsuitable. || 10. A method for use by a system having a hardware processor and a memory storing a software code, the method comprising: receiving, by the software code executed by the hardware processor, dialogue data, the dialogue data identifying a character, a storyline including the character, and a speech for the character intended to at least one of advance the storyline or achieve a goal of the speech; assessing, by the software code executed by the hardware processor and using the dialogue data, a plurality of quality assurance (QA) metrics of the speech, the plurality of QA metrics including at least one of: (i) a fluency of the speech, (ii) a responsiveness of the speech to speech by an interaction partner of the character, (iii) a consistency of the speech with the goal of the speech, (iv) a consistency of the speech with a character profile of the character, or (v) a consistency of the speech with a story-world of the storyline; determining, by the software code executed by the hardware processor and using the plurality of QA metrics, whether the speech is suitable for advancing the storyline; when determining determines that the speech is suitable for advancing the storyline or achieving the goal, approving, by the software code executed by the hardware processor, the speech; and when determining determines that the speech is unsuitable for advancing the storyline or achieving the goal, flagging, by the software code executed by the hardware processor, the speech as being unsuitable. || 19. A system comprising: a hardware processor; and a memory storing a software code; the hardware processor configured to execute the software code to: receive dialogue data, the dialogue data identifying a character, a storyline including the character, and a speech for the character intended to at least one of advance the storyline or achieve a goal of the speech, the speech including a plurality of alternative lines of dialogue; assess, using the dialogue data, a plurality of quality assurance (QA) metrics of the speech, the plurality of QA metrics including at least one of: (i) a fluency of the speech, (ii) a responsiveness of the speech to speech by an interaction partner of the character, (iii) a consistency of the speech with the goal of the speech, (iv) a consistency of the speech with a character profile of the character, or (v) a consistency of the speech with a story-world of the storyline; and determine, using the plurality of QA metrics, one of the alternative lines of dialogue as a best speech to advance the storyline or achieve the goal.