Outer Rim Archives
Archives · 2021 · 10992902

Granted patent

Aspect ratio conversion with machine learning

Number
10992902
Published
2021-04-27
Filed
2019-03-21
Assignee
Disney Enterprises, Inc.
Inventors
Narayan; Nimesh C., Yaacob; Yazmaliza, Grubin; Kari M., Wahlquist; Andrew J.
CPC
H04N21/234372; G06N3/0464; G06N3/084; G06F18/41; G06V40/20; H04N21/44008; G06N5/025; H04N21/2335; H04N21/466; G06F3/04845; G06N3/09; H04N7/0122; G06V40/168; G06F18/211; G06V20/35; G06V40/172; H04N21/4398; G06V20/41; H04N21/440272; G06V20/46
Verdict
Set aside aspect ratio conversion, generic video processing
Source
Google Patents · FreePatentsOnline

Abstract

Techniques are disclosed for converting image frames, such as the image frames of a motion picture, from one aspect ratio to another while predicting the pan and scan framing decisions that a human operator would make. In one configuration, one or more functions for predicting pan and scan framing decisions are determined, at least in part, via machine learning using training data that includes historical pan and scan conversions. The training data may be prepared by extracting features indicating visual and/or audio elements associated with particular shots, among other things. Function(s) may be determined, using machine learning, that take such extracted features as input and output predicted pan and scan framing decisions. Thereafter, the image frames of a received video may be converted between aspect ratios on a shot-by-shot basis, by extracting the same features and using the function(s) to make pan and scan framing predictions.

Background

BACKGROUND Field of the Disclosure (1) Aspects presented in this disclosure generally relate to converting image frames of a video between aspect ratios. Description of the Related Art (2) The aspect ratio of an image is a proportional relationship between a width of the image and a height of the image. Motion picture content is typically created for theatrical distribution in a wide-screen aspect ratio, such as the 2.39:1 aspect ratio, that is wider than aspect ratios supported by in-home display devices, such as the 1.78:1 aspect ratio for high-definition televisions. As a result, image frames of motion pictures in wide-screen aspect ratios need to be converted to smaller aspect ratios for home video distribution. One approach for aspect ratio conversion is pan and scan, which entails reframing image frames to preserve the primary focus of scenes depicted therein, while cropping out the rest of the image frames (e.g., the sides of wide-screen aspect ratio image frames). The faming decisions may include shifting the preserved window left or right to follow the action in a scene, creating the effect of a “pan” shot, and ensuring that the focus of each shot is preserved (i.e., “scanned”). Traditional pan and scan requires manual decisions to be made on how to represent on-screen action throughout a motion picture in a new aspect ratio, which can be both labor intensive and time consuming. SUMMARY (3) One aspect of this disclosure provides a computer-implemented method for conv

Claims

1. A computer-implemented method, the method comprising: determining, using a machine learning (ML) model, an aspect ratio conversion function configured to convert images from a first aspect ratio to a second aspect ratio, wherein the ML model is trained using both (i) first extracted features which indicate at least one of visual or audio elements associated with a first plurality of image frames in shots from videos in the first aspect ratio, and (ii) manual framing decisions in pan and scan conversions of the first plurality of image frames in the first aspect ratio to a corresponding second plurality of image frames in the second aspect ratio; identifying shots in a received video, wherein the shots comprise a third plurality of image frames in the first aspect ratio; and for each of the third plurality of image frames in the identified shots, converting the respective image frame from the first aspect ratio to the second aspect ratio using the determined aspect ratio conversion function. || 11. A non-transitory computer-readable storage medium including instructions that, when executed by a processing unit, cause the processing unit to perform operations comprising: determining, using a machine learning (ML) model, an aspect ratio conversion function configured to convert images from a first aspect ratio to a second aspect ratio, wherein the ML model is trained using both (i) first extracted features which indicate at least one of visual or audio elements associated with a first plurality of image frames in shots from videos in the first aspect ratio, and (ii) manual framing decisions in pan and scan conversions of the first plurality of image frames in the first aspect ratio to a corresponding second plurality of image frames in the second aspect ratio; identifying shots in a received video, wherein the shots comprise a third plurality of image frames in the first aspect ratio; and for each of the third plurality of image frames in the identified shots, converting the respective image frame from the first aspect ratio to the second aspect ratio using the determined aspect ratio conversion function. || 20. A system, comprising: one or more processors; and a memory containing a program that, when executed on the one or more processors, performs operations comprising: determining, using a machine learning (ML) model, an aspect ratio conversion function configured to convert images from a first aspect ratio to a second aspect ratio, wherein the ML model is trained using both (i) first extracted features which indicate at least one of visual or audio elements associated with a first plurality of image frames in shots from videos in the first aspect ratio, and (ii) manual framing decisions in pan and scan conversions of the first plurality of image frames in the first aspect ratio to a corresponding second plurality of image frames in the second aspect ratio; identifying shots in a received video, wherein the shots comprise a third plurality of image frames in the first aspect ratio; and for each of the third plurality of image frames in the identified shots, converting the respective image frame from the first aspect ratio to the second aspect ratio using the determined aspect ratio conversion function. || 21. A computer-implemented method, the method comprising: determining, using a machine learning (ML) model, an aspect ratio conversion function configured to convert images from a first aspect ratio to a second aspect ratio, wherein the ML model is trained using both (i) first extracted features which indicate at least one of visual or audio elements associated with a first plurality of image frames in shots from videos in the first aspect ratio, and (ii) manual framing decisions in pan and scan conversions of the first plurality of image frames in the first aspect ratio to a corresponding second plurality of image frames in the second aspect ratio; identifying shots in a received video, wherein the shots comprise a third plurality of image frames in the first aspect ratio; identifying that one or more characters in the received video are speaking based, at least in part, on at least one of (i) mouth movement of the one or more characters in the third plurality of image frames, or (ii) recognition of one or more voices and matching of the recognized one or more voices to visual depictions of the one or more characters in the third plurality of image frames; and for each of the third plurality of image frames in the identified, converting the respective image frame from the first aspect ratio to the second aspect ratio using the determined aspect ratio conversion function, comprising: inputting second extracted features indicating second audio elements associated with the third plurality of image frames into the determined aspect ratio conversion function, wherein the second audio elements relate to speech of the one or more characters who are identified as speaking.