Outer Rim Archives
Archives · 2026 · 20260236215

Application (pre-grant publication)

SYSTEM AND METHOD TO PROVIDE PERSONALIZED AUDIO STREAMING AND RENDERING

Number
20260236215
Published
2026-08-13
Filed
2025-04-17
Assignee
Disney Enterprises, Inc.
Inventors
Briand; Manuel
CPC
G06F3/165; G10L21/0364; H03G3/32; H03G9/005; H03G9/025; H04N21/4394; H04N21/4852
Verdict
Set aside streaming, codec, cdn/infra
In edition
2026-W36
Source
Google Patents · FreePatentsOnline

The keeper's note

Systems and methods to provide personalized audio streaming and rendering include receiving a cinematic audio track for selected content and a maximum accessible audio track for the content, and adjustably combining the…

Abstract

Systems and methods to provide personalized audio streaming and rendering include receiving a cinematic audio track for selected content and a maximum accessible audio track for the content, and adjustably combining the cinematic audio track and accessible audio track to provide an improved dialogue audio track which allows a user/listener to hear the dialogue over an environmental noise floor, the combining being provided by a cross-fade renderer adjusted manually by the user/listener or automatically adjusted based on a measured noise floor, and optionally providing a personalized equalizer for hearing impairments. It also allows a content provider to create the maximum accessible audio track for a given content separate from the user/listener, which may also use a X-Fade renderer to verify quality of the accessible audio track.

Background

BACKGROUND

Current audio streaming solutions typically rely on a single audio bitstream and decoder per streaming application which limits the ability of the content owners and streaming service providers to personalize the audio experience it provides to the end user.

In particular, existing steaming services provide separate streams for different versions enhanced dialogue, e.g., English dialogue boost high, English dialogue boost medium, and the like. In that case, the user must select which version they want to listen to. Also, to change to another audio version, e.g., because the background environment noise changed, the user must make a request, and a new version is retrieved from the appropriate content server. Additionally, most TV devices apply post-processing to the decoded audio bitstream, such as AI sound enhancements, to reduce noise or boost dialogue which does not preserve the integrity of the original sound mix nor the creative intent of the content owners.

Switching audio streams can be cumbersome, slow and incur digital streaming file buffering/loading delays during the audio delivery from the Content Delivery Network (CDN) to the end user. Also, conventional systems only provide a pre-defined number of dialogue enhanced versions (e.g., 3-5 versions), without any guarantee that a given selected version may be ideal for the noise environment of the listener/user. Such an approach can make it difficult for the user to hear the dialogue of

Claims

1. A computer-based method for generating a personalized accessible audio mix (Acc Mix) of audio content having dialogue audio and non-dialogue audio that allows at least one user to hear and understand the dialogue audio when played on a user playing device in a listening environment having an actual noise floor (NF) level, comprising: receiving, at the user playing device, an original cinematic audio mix (CinMix) having the dialogue audio and the non-dialogue audio as set from a recording studio; receiving, at the user playing device, a maximum accessible mix (MaxAccMix) of the dialogue audio and the non-dialogue audio, the Max-Acc-Mix being a audio modification of the CinMix such that a loudness range of the dialogue audio and a loudness range of the non-dialogue audio are greater than a predetermined acceptable noise floor and less than a predetermined maximum audio loudness limit; and combining the CinMix and MaxAccMix to create the personalized AccMix audio mix which is power preserving and has an loudness of the dialogue audio that allows a user to hear and understand the dialogue audio when played on the user playing device in a listening environment having the actual noise floor (NF) level, wherein the value of the personalized AccMix ranges from the CinMix value to the MaxAccMix value, and is a power preserving combination of CinMix and MaxAccMix therebetween, and the personalized AccMix value is set based on at least one of user input and the actual noise floor level. || 10. A computer-based method for creating a maximum accessible audio mix (Max-Acc-Mix), comprising: receiving original cinematic mix (Cin-Mix) audio content; and adjusting the Cin-Mix using a dialogue clarity engine (DCE) having predetermined DCE gains to create the Max-Acc-Mix. || 19. A computer-based method for generating an accessible audio mix (Acc Mix) of dialogue audio and non-dialogue audio (residual music/effects) that allows a listener user to understand the dialogue audio played on a user playing device (or listening device) in a listening environment having plurality of different actual noise floor levels, comprising: receiving, at a user device, a cinematic mix (CinMix) and a maximum accessible mix (MaxAccMix) associated with video/audio content to be viewed and listened to by a user from the user device; creating the AccMix audio mix by blending the CinMix and the MaxAccMix such that the resultant AccMix audio mix is power preserving and is greater than a noise floor (NF) in the listening environment of the user device, the AccMix having a range from the CinMix value to the MaxAccMix value, the value of AccMix being based on at least one of user input and the actual noise floor level. || 21. A computer-based method of increasing the audio dialogue level of digital cinematic content while preserving creative intent of the cinematic content, comprising: receiving a cinematic audio mix; receiving an accessible audio mix; and combining the cinematic audio mix and accessible audio mix to provide an improved dialogue audio mix which allows a listener to hear the dialogue over a noise floor. || 23. A computer-based method of fading from a first audio mix (MixA) to a second audio mix (MixB), comprising: receiving the MixA audio mix; receiving the MixB audio mix; combining MixA and MixB using a cross-fade renderer having an output that transitions from MixA to MixB based on the position of a X-Fade slider.