Outer Rim Archives
Archives · 2018 · 20180061057

Application (pre-grant publication)

Video Object Tagging using Synthetic Images and Segmentation Hierarchies

Number
20180061057
Published
2018-03-01
Filed
2016-08-23
Assignee
DISNEY ENTERPRISES, INC.
Inventors
FARRE GUIU; MIQUEL ANGEL et al.
CPC
G06V20/70; G06F16/316; G06F16/7837; G06F16/7867; G06F16/81; G06T7/11; G06T7/12; G06T7/215; G06T11/60; G06V10/993; G06V20/41
Verdict
Set aside synthetic-image video-object tagging, content indexing
Source
Google Patents · FreePatentsOnline

Abstract

There is provided a system including a memory and a processor configured to obtain a first frame of a video content including an object and a first region based on a segmentation hierarchy of the first frame, insert a synthetic object into the first frame, merge an object segmentation hierarchy of the synthetic object with the segmentation hierarchy of the first frame to create a merged segmentation hierarchy, select a second region based on the merged segmentation hierarchy, provide the first frame including the first region and the second region to a crowd user for creating a corrected frame, receive the corrected frame from the crowd user including a first corrected region including the object and a second corrected region including the synthetic object, determine a quality based on the synthetic object and the second corrected region, and accept the first corrected region based on the quality.

Background

BACKGROUND

Image and video segmentation is one of the most fundamental yet challenging problems in computer vision. Dividing an image into meaningful regions requires a high level interpretation of the image that cannot be satisfactorily solved by only looking for homogeneous areas in an image. In the era of big data and vast computing power, one approach to model high level interpretation of images has been to use powerful machine-learning tools on huge annotated databases. While significant advances have been made in recent years, automatic image segmentation is still far from providing accurate results in a generic scenario. The creator of a video may desire to add information or a link to an object in a video, and may wish the added information or link to remain associated with that object throughout a video sequence.SUMMARY

The present disclosure is directed to video object tagging using synthetic objects and segmentation hierarchies, substantially as shown in and/or described in connection with at least one of the figures, as set forth more completely in the claims.

Claims

1. A system for tagging an object in a video content including a plurality of frames, the system comprising: a memory storing an executable code; and a processor executing the executable code to: obtain a first frame of the video content including the object and a first region including at least part of the object based on a segmentation hierarchy of the first frame; insert a synthetic object into the first frame of the video content; merge an object segmentation hierarchy of the synthetic object with the segmentation hierarchy of the first frame to create a merged segmentation hierarchy; select a second region at least partially including the synthetic object based on the merged segmentation hierarchy; provide the first frame including the first region and the second region to a first crowd user for creating a first corrected frame; receive the first corrected frame from the first crowd user including a first corrected region including the object and a second corrected region including the synthetic object; determine a quality based on the synthetic object and the second corrected region; and accept the first corrected region based on the quality. 11. A method for tagging an object in a video content including a plurality of frames, for use with a system having a non-transitory memory and a hardware processor, the method comprising: obtaining, using the hardware processor, a first frame of the video content including the object and a first region including at least part of the object based on a segmentation hierarchy of the first frame; inserting, using the hardware processor, a synthetic object into the first frame of the video content; merging, using the hardware processor, an object segmentation hierarchy of the synthetic object with the segmentation hierarchy of the first frame to create a merged segmentation hierarchy; selecting, using the hardware processor, a second region at least partially including the synthetic object based on the merged segmentation hierarchy; providing, using the hardware processor, the first frame including the first region and the second region to a first crowd user for creating a first corrected frame; receiving, using the hardware processor, the first corrected frame from the first crowd user including a first corrected region including the object and a second corrected region including the synthetic object; determining, using the hardware processor, a quality based on the synthetic object and the second corrected region; and accepting, using the hardware processor, the first corrected region based on the quality.