Outer Rim Archives
Archives · 2022 · 20220309345

Application (pre-grant publication)

Machine Learning Model Based Embedding for Adaptable Content Evaluation

Number
20220309345
Published
2022-09-29
Filed
2022-03-16
Assignee
Disney Enterprises, Inc.
Inventors
Doggett; Erika Varis, Beard; Audrey Coyote Aura, Schroers; Christopher Richard, Azevedo; Roberto Gerson de Albuquerque, Labrozzi; Scott, Xue; Yuanyi, Zimmerman; James
CPC
G06N3/0895; G06N3/0464; G06N3/045; G06N3/08
Verdict
Set aside content evaluation embedding, business
Source
Google Patents · FreePatentsOnline

Abstract

A system includes a computing platform having processing hardware, and a system memory storing software code and one or more machine learning (ML) model(s) trained using contrastive learning based on a similarity metric. The processing hardware is configured to execute the software code to receive input data including a plurality of content segments, map, using the ML model(s), each of the plurality of content segments to a respective embedding in a continuous vector space to provide a plurality of mapped embeddings, and perform one of a classification or a regression of the content segments using the plurality of mapped embeddings. The processing hardware is also configured to execute the software code to discover, based on the classification or the regression, at least one new label for characterizing the plurality of content segments.

Background

BACKGROUND

Due to its nearly universal popularity as a content medium, ever more visual media content is being produced and grade available to consumers. As a result, the efficiency with which visual images can be analyzed, classified, and processed has become increasingly important to the producers, owners, and distributors of that visual media content.

One significant challenge to the efficient classification and processing of visual media content is that entertainment and media studios produce many different types of content having differing features, such as different visual textures and movement. In the case of audio-video (AV) film and television content, for example the content produced may include live action content with realistic computer-generated imagery (CGI) elements, high complexity three-dimensional (3D) animation, and even two-dimensional (2D) hand-drawn animation. Moreover, each different type of content produced may require different treatment in pre-production, post-production, or both.

Consider, for example, the post-production treatment of AV or video content. Different types of AV or video content may benefit from different encoding schemes for streaming, or different workflows for localization. In the conventional art, the classification of content as being of a particular type is typically done manually, through human inspection, and in the example use case of video encoding, the most appropriate workflow may not be identifiable e

Claims

1. A system comprising: a processing hardware; and a system memory storing a software code and at least one machine learning (ML) model trained using contrastive learning based on a similarity metric; the processing hardware configured to execute the software code to: receive an input including a plurality of content segments; map, using the at least one ML model, each of the plurality of content segments to a respective embedding in a continuous vector space to provide a plurality of mapped embeddings corresponding respectively to the plurality of content segments; perform one of a classification or a regression of the content segments using the plurality of mapped embeddings; and discover, based on the ne f the class on or the regression, at least one new label for characterizing the plurality of content segments. || 11. A method for use by a system including a processing hardware, and a system memory storing a software code and at least one machine learning (ML) model trained using contrastive learning based on a similarity metric, the method comprising: receiving, by the software code executed by the processing hardware, an input including a plurality of content segments; mapping, by the software code executed by the processing hardware and using the at least one ML model, each of the plurality of content segments to a respective embedding in a continuous vector space to provide a plurality of mapped embeddings corresponding respectively to the plurality of content segments; performing one of a classification or a regression of the content segments, by the software code executed by the processing hardware, using the plurality of mapped embeddings; and discovering, by the software code executed by the processing hardware based on the one of the classification or the regression, at least one new label for character the plurality of content segments.