- Number
- 20250335817
- Published
- 2025-10-30
- Filed
- 2024-04-30
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Tiwari; Sanchita
- CPC
- G06N20/00
- Verdict
- Low Notable software
- Source
- Google Patents · FreePatentsOnline
The keeper's note
ML-based interactive digital-persona generation technique (Tiwari, character AI cluster).
Abstract
A system includes a hardware processor configured to execute a machine learning (ML) model training pipeline to train an ML model using data relevant to a world of a digital persona to provide a dialogue model, generate, using the dialogue model, first conversational outputs, train the dialogue model, based on the first conversational outputs, to avoid hallucinations and/or undesirable expressions to provide a guardrailed dialogue model, generate, using the guardrailed dialogue model, second conversational outputs, train the guardrailed dialogue model, based on the second conversational outputs and persona data identifying interaction characteristics of the digital persona to provide a persona-specific model, generate, using the persona-specific model, a response to a scripted question, determine a quality score for the response, and further train the persona-specific model or validate the persona-specific model for human interaction, depending upon whether the quality score fails to satisfy or satisfies a quality criterion.
Background
BACKGROUND
Advances in artificial intelligence (AI) have led to the development of systems capable of interacting with a human user in a variety of ways. Large language models, for example, have shown tremendous potential in creating complex and nuanced conversations with users. However, existing large language models typically present themselves as generic interaction portals lacking a distinctive personality. As a result, and although large language models are generally successful in holding a conversation or providing requested information, they fail to project the type of persona that can encourage a user to develop an emotional connection or affinity with the persona model. Thus, imbuing a large language model with the semblance of a personality may improve the user experience and give rise to a sense of loyalty or even affection for a particular persona model.
Nevertheless, training a large language model to project a consistent personality that is substantially guard railed against hallucination and the generation of toxic, offensive, or otherwise undesirable language presents significant challenges. For instance, large language models may include well over one hundred billion parameters, and require the expenditure of enormous resource to guardrail and train. Consequently, there is a need in the art for an efficient and resource sparing solution for imbuing machine learning models such as large language models with distinctive digital personas.
Claims
1. A system comprising: a hardware processor, and a memory storing a machine learning (ML) model training pipeline; the hardware processor configured to execute the ML model training pipeline to: train an ML model using a dataset including data relevant to a world of a predetermined digital persona, to provide a dialogue model; generate, using the dialogue model, a plurality of first conversational outputs; train the dialogue model, based on the plurality of first conversational outputs, to avoid at least one of hallucinations or undesirable expressions, to provide a guardrailed dialogue model; generate, using the guardrailed dialogue model, a plurality of second conversational outputs; train the guardrailed dialogue model, based on the plurality of second conversational outputs and persona data identifying interaction characteristics of the predetermined digital persona, to provide a persona-specific model; generate, using the persona-specific model, a response to a scripted question; determine a quality score for the response; further train the persona-specific model, when the quality score fails to satisfy a quality criterion, or validate the persona-specific model for human interaction when the quality score satisfies the quality criterion. ||
11. A method for use by a system including a hardware processor, and a memory storing a machine learning (ML) model training pipeline, the method comprising: training an ML model, by the hardware processor using the ML model training pipeline and a dataset including data relevant to a world of a predetermined digital persona, to provide a dialogue model; generating, by the hardware processor using the ML model training pipeline and the dialogue model, a plurality of first conversational outputs; training the dialogue model, by the hardware processor using the ML model training pipeline based on the plurality of first conversational outputs, to avoid at least one of hallucinations or undesirable expressions, to provide a guardrailed dialogue model; generating, by the hardware processor using the ML model training pipeline and the guardrailed dialogue model, a plurality of second conversational outputs; training the guardrailed dialogue model, by the hardware processor using the ML model training pipeline based on the plurality of second conversational outputs and persona data identifying interaction characteristics of the predetermined digital persona, to provide a persona-specific model; generating, by the hardware processor using the ML model training pipeline and using the persona-specific model, a response to a scripted question; determining, by the hardware processor using the ML model training pipeline, a quality score for the response; further training the persona-specific model, by the hardware processor using the ML model training pipeline when the quality score fails to satisfy a quality criterion, or validating the persona-specific model for human interaction by the hardware processor using the ML model training pipeline when the quality score satisfies the quality criterion.