Outer Rim Archives
Archives · 2026 · 20260268939

Application (pre-grant publication)

SYSTEM AND METHOD FOR VIDEO/AUDIO COMPREHENSION AND AUTOMATED CLIPPING

Number
20260268939
Published
2026-09-10
Filed
2026-04-30
Assignee
Disney Enterprises, Inc.
Inventors
Hyde; Andrew, Booth; Geoffrey, Boras; Ognjen, Flanders; Jonathan, Donnell; Danny
CPC
G11B27/34; G06F40/205; G06F40/295; H04N21/8549
Verdict
Set aside nlp/localization
In edition
2026-W37
Source
Google Patents · FreePatentsOnline

The keeper's note

Systems and Methods for Video/Audio Comprehension and Automated Clipping includes providing at least one media clip (MC) within an event for display or listening on a user device including receiving audio or video media…

Abstract

Systems and Methods for Video/Audio Comprehension and Automated Clipping includes providing at least one media clip (MC) within an event for display or listening on a user device including receiving audio or video media data indicative of the event, transcribing the media data into timestamped text, identifying entities within the text, creating text segments having a begin timestamp and end timestamp and having a minimum number of entity mentions in the text segments, clipping from the media data the at least one media clip having a begin timestamp and end timestamp corresponding to the begin timestamp and end timestamp of a corresponding one of the text segments, and providing the at least one media clip to the user device for viewing or listening by a user. Feedback may also be provided to adjust the logic that identifies MCs. MC Alerts may also be sent to users autonomously or based on user-set parameters.

Background

BACKGROUND

The process of reviewing and analyzing long-form and live media content, such as video and audio, to create short-form video-on-demand (VOD) or audio-on-demand (AOD) clips or segments for consumption by users, is a process that requires extensive time due to manual processes. Such media clips are currently created via manual inspection and detailed review of the media content to identify specific topics in the longer video, e.g., a sports show or event or other shows or events, to create the clips. Such a process is slow and results in a very limited quantity of short VOD media clips for consumption by users or consumers.

Accordingly, it would be desirable to have a system and method that increases the amount of such clips and decreases the time to create them, thereby providing a greater number of short VOD/AOD media content clips that are of interest to sports fans or the general public.

Claims

1. An automated computer-based system for providing at least one media clip (MC) from an event for display or listening on a user device, comprising: a processor configured to receive media data indicative of the event; the processor further configured to transcribe the audio portion of the media data into text with timestamps; the processor further configured to identify entities within the text, the entities being named in the content of the text; the processor further configured to perform phonetic correction and co-reference resolution of the entities using predetermined phonetic rules and predetermined co-reference rules, respectively; the processor further configured to segment the text into a plurality of text segments based on predetermined text segment creation rules, each of the text segments having at least one of the entities and having a segment begin timestamp and a segment end timestamp; the processor further configured to clip from the media data the at least one media clip having a clip begin timestamp and clip end timestamp that corresponds to the segment begin timestamp and the segment end timestamp of a corresponding one of the text segments; and the processor further configured to provide the at least one media clip for viewing or listening on the user device, wherein the processor being further configured to identify, perform, segment, and clip contiguously in an automated manner without human intervention after having received the media data, using the predetermined phonetic rules, the predetermined co-reference rules, and the predetermined text segment creation rules. || 22. An automated computer-based system for providing at least one media clip (MC) from an event for display on a user device, comprising: a processor configured to receive media data indicative of the event, the media data having video timestamps; the processor further configured to transcribe the audio portion of the media data into timestamped text; the processor further configured to identify one or more entities within the text, the entities being named in the content of the text; the processor further configured to perform phonetic correction and co-reference resolution of the entities; the processor further configured to segment the text into a plurality of text segments, each of the text segments having at least one of the entities and having a segment begin timestamp and a segment end timestamp; the processor further configured to determine an entity classification of the at least one entity comprising an amount of time the at least one entity is mentioned during a given segment or during the entire event; the processor further configured to clip from the media data the at least one media clip having a media clip begin timestamp and media clip end timestamp that corresponds to the segment begin timestamp and the segment end timestamp of a corresponding one of the text segments; and the processor further configured to provide the at least one media clip with the entity classification to the user device, the user device being configured to show the at least one media clip, wherein the processor is further configured to identify, perform, segment, determine, and clip contiguously in an automated manner without human intervention after receiving the media data, using predetermined rules. || 24. An automated computer-based system for providing at least one media clip (MC) from a sports event for display on a user device, comprising: a processor configured to receive media data indicative of the event, the media data having an audio channel and a video channel, the video channel having video timestamps; the processor further configured to transcribe the audio channel portion of the media data into text with timestamps; the processor further configured to identify entities being named in the content of the text and associate the entities in the content of the text with at least one of corresponding pronouns, relationship words, nicknames, and abbreviations; the processor further configured to create a plurality of text segments from the text, each of the text segments having the at least one of the entities and having a segment begin timestamp and a segment end timestamp; the processor further configured to extract from the media data the at least one media clip having a media clip begin timestamp and a media clip end timestamp that corresponds to the segment begin timestamp and segment end timestamp of a corresponding one of the text segments; and the processor further configured to provide the at least one media clip to the user device for viewing by a user, wherein the processor being further configured to identify, create, and extract contiguously in an automated manner without human intervention after having received the media data, using predetermined rules. || 25. An automated computer-based system for providing at least one media clip (MC) from an event comprising: a processor configured to receive media data indicative of the event; the processor further configured to transcribe the media data into timestamped text; the processor further configured to identify entities within the text, the entities being named in the content of the text; the processor further configured to create text segments having at least one of the entities and having a segment begin timestamp and a segment end timestamp and having a minimum number of entity mentions in the text segments within a maximum entity gap length time; the processor further configured to clipp from the media data the at least one media clip having a clip begin timestamp and a clip end timestamp corresponding to the segment begin timestamp and the segment end timestamp of a corresponding one of the text segments; and wherein the processor being further configured to identify, create, and clip contiguously in an automated manner without human intervention after receiving the media data, using predetermined rules. || 35. An automated computer-based system for identifying and classifying entities in text and providing entity classification tagged text, comprising: a processor configured to receive text data, which is a transcription with timestamps of an audio portion of media data; the processor further configured to identify at least one entity within the text data, the at least one entity being named in the content of the text; the processor further configured to segment the text into a plurality of text segments based on predetermined text segment creation rules, each of the text segments having at least one of the entities and having a segment; the processor further configured to determine an entity classification of the at least one entity comprising an amount that the at least one entity is mentioned during a given segment or during the entire event, and wherein the text segment includes the entity classification; and wherein the processor being further configured to identify, perform, segment, and determine contiguously in an automated manner without human intervention after receiving the text data, using the predetermined rules.