Outer Rim Archives
Archives · 2024 · 12003831

Granted patent

Automated content segmentation and identification of fungible content

Number
12003831
Published
2024-06-04
Filed
2021-07-02
Assignee
Disney Enterprises, Inc.
Inventors
Farre Guiu; Miquel Angel et al.
CPC
H04N21/8352; H04N21/251; H04N21/233; G06N20/00; H04N21/23418; H04N21/4662; G06V20/41; H04N21/8456
Verdict
Set aside content segmentation/fungibility classification, business
Source
Google Patents · FreePatentsOnline

Abstract

A content segmentation system includes a computing platform having processing hardware and a system memory storing a software code and a trained machine learning model. The processing hardware is configured to execute the software code to receive content, the content including multiple sections each having multiple content blocks in sequence, to select one of the sections for segmentation, and to identify, for each of the content blocks of the selected section, at least one respective representative unit of content. The software code is further executed to generate, using the at least one respective representative unit of content, a respective embedding vector for each of the content blocks of the selected section to provide a multiple embedding vectors, and to predict, using the trained machine learning model and the embedding vectors, subsections of the selected section, at least some of the subsections including more than one of the content blocks.

Background

BACKGROUND (1) Due to its popularity as a content medium, ever more video in the form of episodic television (TV) and movie content is being produced and made available to consumers via streaming services. As a result, the efficiency with which segments of a video content stream having different bit-rate encoding requirements are identified has become increasingly important to the producers and distributors of that video content. (2) Segmentation of video and other content has traditionally been performed manually by human editors. However, such manual segmentation of content is a labor intensive and time consuming process. Consequently, there is a need in the art for an automated solution for performing content segmentation that substantially minimizes the amount of content, such as audio content and video content, requiring manual processing.

Claims

1. A system comprising: a computing platform including processing hardware and a system memory; a software code stored in the system memory; and a machine learning model trained to predict scene boundaries; the processing hardware configured to execute the software code to: receive content, the content including a plurality of acts each having a plurality of shots in sequence; select one of the plurality of acts for segmentation; identify, for each of the plurality of shots of the selected act, at least one respective representative unit of content, the at least one respective unit of content comprising at least one video frame, an audio sample, or a combination thereof; generate, using the at least one respective representative unit of content, a respective embedding vector for each of the plurality of shots of the selected act to provide a plurality of embedding vectors; and predict, using the machine learning model and the plurality of embedding vectors, respective scene boundaries of each of a plurality of scenes of the selected act, at least one of the plurality of scenes including more than one of the plurality of shots. || 11. A computer-readable non-transitory storage medium having stored thereon instructions, which when executed by a processing hardware of a system, instantiate a method comprising: receiving content, the content including a plurality of acts each having a plurality of shots in sequence; selecting one of the plurality of acts for segmentation; identifying, for each of the plurality of shots of the selected act, at least one respective representative unit of content, the at least one respective unit of content comprising at least one video frame, an audio sample, or a combination thereof; generating, using the at least one respective representative unit of content, a respective embedding vector for each of the plurality of shots of the selected act to provide a plurality of embedding vectors; and predicting, using a machine learning model and the plurality of embedding vectors, respective scene boundaries of each of a plurality of scenes of the selected act, at least one of the plurality of scenes including more than one of the plurality of shots. || 17. A method for use by a system including a computing platform having processing hardware, a system memory storing a software code, and a machine learning model trained to predict scene boundaries, the method comprising: receiving content, by the software code executed by the processing hardware, the content including a plurality of acts each having a plurality of shots in sequence; selecting, by the software code executed by the processing hardware, one of the plurality of acts for segmentation; identifying, by the software code executed by the processing hardware for each of the plurality of shots of the selected act, at least one respective representative unit of content, the at least one respective unit of content comprising at least one video frame, an audio sample, or a combination thereof; generating, by the software code executed by the processing hardware and using the at least one respective representative unit of content, a respective embedding vector for each of the plurality of shots of the selected act to provide a plurality of embedding vectors; and predicting, by the software code executed by the processing hardware using the machine learning model and the plurality of embedding vectors, respective scene boundaries of each of a plurality of scenes of the selected act, at least one of the plurality of scenes including more than one of the plurality of shots.