Outer Rim Archives
Archives · 2026 · 20260245545

Application (pre-grant publication)

INTERACTIVE IMMERSIVE EXPERIENCES USING MACHINE LEARNING

Number
20260245545
Published
2026-08-20
Filed
2025-02-18
Assignee
Lucasfilm Entertainment Company Ltd. LLC
Inventors
Peck; Alexander, Koperwas; Michael
CPC
G10L13/10; G10L13/027; G10L15/183; G10L15/22
Verdict
Medium Notable software
First reported
2026-W36 (2026-09-05)
Source
Google Patents · FreePatentsOnline

The keeper's note

Machine learning that tunes an interactive immersive experience to how guests are engaging with it.

Abstract

Methods and systems for improving interactive immersive experiences using machine learning are disclosed. In an interactive media experience, a user may speak to virtual characters displayed to the user (e.g., via a screen on a user device). The user's speech may be recorded to produce user speech data, which can be processed using a response model (e.g., a large language model) to generate a textual response and one or more tonal indicators. The one or more tonal indicators can be used to identify text-to-speech models that can be used to generate an audio response based on the textual response, which can then be played back to the user (e.g., via a speaker on the user device), thereby effectively communicating tonal information and achieving a more immersive user experience. Various other improvements to interactive media experiences are also disclosed, including methods for reducing communication latency and dealing with speech interruptions.

Background

BACKGROUND

Many forms of story-based entertainment media (e.g., books, movies, television, videogames, etc.) rely on scripted story content, e.g., stories that have been prepared in advanced by writers. Such stories are often fixed, unchanging, and have limited interactivity. For example, it is rare for a movie to have multiple endings, and while videogames may offer a player some influence over a story (e.g., by playing through the “hero story mode” or the “villain story mode”), usually such influence is limited, e.g., by letting the player select one of three options in a dialog tree. As a result, such media can often have limited “replay value”. This can be particularly noticeable in alternative entertainment forms, such as augmented reality (AR) or virtual reality (VR) media in which a user is “immersed” in the story itself, e.g., by taking on the role of a character in the AR or VR environment. While a reader can typically appreciate that the story in a book is unchanging and outside the reader's control, a user in an AR environment typically expects characters and the environment to react to the user in generally logic ways. When characters and the environment do not react to the user appropriately, it can be an “immersion breaking” experience and can negatively impact the user's enjoyment of the media.

Some forms of story-based entertainment are improvisational or reactive. Such entertainment usually relies on human entertainers or some form of media oper

Claims

1. A method performed by a computer system for generating an audio response, the method comprising: receiving user speech data corresponding to a user; generating, using a response model, a textual response to the user speech data and one or more tonal indicators based on the user speech data; determining, based on the one or more tonal indicators, one or more text-to-speech models corresponding to the one or more tonal indicators; generating, using the one or more text-to-speech models, an audio response to the user speech data; and causing the audio response to be played to the user. || 19. A method performed by a user device for generating an audio response and playing the audio response to a user, the method comprising: recording the user, thereby generating a user audio recording; generating, based on the user audio recording, user speech data; generating, using a response model, a textual response and one or more tonal indicators corresponding to the user speech data; determining, based on the one or more tonal indicators, one or more text-to-speech models corresponding to the one or more tonal indicators; generating, using the one or more text-to-speech models, the audio response to the user speech data; and playing the audio response to the user.