Outer Rim Archives
Archives · 2026 · 20260244679

Application (pre-grant publication)

MULTIMODAL MACHINE LEARNING MODEL FOR CONTENT EVALUATION

Number
20260244679
Published
2026-08-20
Filed
2026-04-14
Assignee
Disney Enterprises, Inc.
Inventors
GAO; Yupeng, GAO; Pengfei, ZHANG; Yan, WANG; Zhe, LI; Mengzhe, XIAO; Xingpeng, HOSSAIN; Yasir, MILANO; Gianluca
CPC
G06F16/435; G06F16/24578; G06N3/044; G06N3/045; G06N3/0464; G06N3/047; G06N3/08; G06N3/084; G06N3/09; G06N7/01; G06N20/00; G06N20/10; G06N20/20; G06F16/24578
Verdict
Set aside streaming, recommendation
In edition
2026-W36
Source
Google Patents · FreePatentsOnline

The keeper's note

Embodiments provide for improved machine learning.

Abstract

Embodiments provide for improved machine learning. A request for supplemental content to be provided in association with a media content item is received, and a set of candidate supplemental content items for the request is determined. A user embedding corresponding to a user embedding corresponding to a user associated with the media content item, a media embedding corresponding to the media content item, and a set of supplemental content embeddings corresponding to the set of candidate supplemental content items are accessed from one or more storage repositories. A set of interaction scores is generated based on processing the user embedding, the media embedding, and the set of supplemental content embeddings using an interaction machine learning model. A first supplemental content item of the set of candidate supplemental content items is selected for the request based on the set of interaction scores.

Background

BACKGROUND

The digital content landscape is continuously evolving. Not only is there a tremendous variety of primary content (e.g., multimedia such as a video stream, audio stream, and the like) available to users, but there is also a similarly vast assortment of supplemental content (e.g., promotional content, recommendations, live events, and the like) which can be provided along with the primary content. Though significant resources have been expended seeking to improve supplemental content selection, there remains substantial opportunity for improvement. Recently, some attempts have been made to use machine learning to improve content selection. However, such approaches have thus far been suboptimal in their selections. Further, such approaches generally incur substantial computational expense (e.g., relying on substantial compute resources such as memory). Further, such approaches generally introduce significant latency (e.g., significant time is consumed processing the various data to select content), rendering these approaches unsuitable for many digital content environments where these delays are unacceptable.

Claims

1. A method, comprising: receiving a request for supplemental content to be provided in association with a media content item; determining a set of candidate supplemental content items for the request; accessing, from one or more storage repositories, a user embedding corresponding to a user associated with the media content item, a media embedding corresponding to the media content item, and a set of supplemental content embeddings corresponding to the set of candidate supplemental content items; generating a set of interaction scores based on processing the user embedding, the media embedding, and the set of supplemental content embeddings using an interaction machine learning model, wherein the set of interaction scores indicate predicted cohesion between the media content item and each of the candidate supplemental content items; and selecting, for the request, a first supplemental content item of the set of candidate supplemental content items based on the set of interaction scores. || 11. One or more non-transitory computer readable media containing, in any combination, computer program code that, when executed by operation of any combination of one or more processors, performs an operation comprising: receiving a request for supplemental content to be provided in association with a media content item; determining a set of candidate supplemental content items for the request; accessing, from one or more storage repositories, a user embedding corresponding to a user associated with the media content item, a media embedding corresponding to the media content item, and a set of supplemental content embeddings corresponding to the set of candidate supplemental content items; generating a set of interaction scores based on processing the user embedding, the media embedding, and the set of supplemental content embeddings using an interaction machine learning model, wherein the set of interaction scores indicate predicted cohesion between the media content item and each of the candidate supplemental content items; and selecting, for the request, a first supplemental content item of the set of candidate supplemental content items based on the set of interaction scores. || 17. A system, comprising: one or more processors; and one or more memories storing a program, which, when executed on any combination of the one or more processors, performs operations, the operations comprising: receiving a request for supplemental content to be provided in association with a media content item; determining a set of candidate supplemental content items for the request; accessing, from one or more storage repositories, a user embedding corresponding to a user associated with the media content item, a media embedding corresponding to the media content item, and a set of supplemental content embeddings corresponding to the set of candidate supplemental content items; generating a set of interaction scores based on processing the user embedding, the media embedding, and the set of supplemental content embeddings using an interaction machine learning model, wherein the set of interaction scores indicate predicted cohesion between the media content item and each of the candidate supplemental content items; and selecting, for the request, a first supplemental content item of the set of candidate supplemental content items based on the set of interaction scores.