- Number
- 10664973
- Published
- 2020-05-26
- Filed
- 2018-07-10
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Farre Guiu; Miquel Angel, Martin; Marc Junyent, Smolic; Aljoscha
- CPC
- G06V20/70; G06F16/316; G06F16/7837; G06F16/7867; G06F16/81; G06T7/11; G06T7/12; G06T7/215; G06T11/60; G06V10/993; G06V20/41
- Verdict
- Set aside video object tagging for content management, business
- Source
- Google Patents · FreePatentsOnline
Abstract
There is provided a system including a memory and a processor configured to obtain a first frame of a video content including an object and a first region based on a segmentation hierarchy of the first frame, insert a synthetic object into the first frame, merge an object segmentation hierarchy of the synthetic object with the segmentation hierarchy of the first frame to create a merged segmentation hierarchy, select a second region based on the merged segmentation hierarchy, provide the first frame including the first region and the second region to a crowd user for creating a corrected frame, receive the corrected frame from the crowd user including a first corrected region including the object and a second corrected region including the synthetic object, determine a quality based on the synthetic object and the second corrected region, and accept the first corrected region based on the quality.
Background
BACKGROUND(1) Image and video segmentation is one of the most fundamental yet challenging problems in computer vision. Dividing an image into meaningful regions requires a high level interpretation of the image that cannot be satisfactorily solved by only looking for homogeneous areas in an image. In the era of big data and vast computing power, one approach to model high level interpretation of images has been to use powerful machine-learning tools on huge annotated databases. While significant advances have been made in recent years, automatic image segmentation is still far from providing accurate results in a generic scenario. The creator of a video may desire to add information or a link to an object in a video, and may wish the added information or link to remain associated with that object throughout a video sequence.SUMMARY(2) The present disclosure is directed to video object tagging using synthetic objects and segmentation hierarchies, substantially as shown in and/or described in connection with at least one of the figures, as set forth more completely in the claims.
Claims
1. A system comprising: a memory storing an executable code; and a processor executing the executable code to: obtain a first frame of a video content including an object and a first region surrounding at least part of the object, the first region having one or more first leak areas; insert a synthetic object obtained from a synthetic object database into the first frame of the video content, the synthetic object having a shape stored in the synthetic object database; select a second region at least partially surrounding the synthetic object, the second region having one or more second leak areas; provide the first frame including the first region and the second region to a first user for creating a first corrected frame; receive the first corrected frame including one or more first corrections by the first user to the one or more first leak areas of the first region, and one or more second corrections by the first user to the one or more second leak areas of the second region; determine a quality of the one or more second corrections to the one or more second leak areas of the second region based on the shape of the synthetic object; and accept the one or more first corrections to the one or more first leak areas of the first region depending on the quality of the one or more second corrections to the one or more second leak areas of the second region.
8. A method for use with a system having a non-transitory memory and a hardware processor, the method comprising: obtaining, using the hardware processor, a first frame of a video content including an object and a first region surrounding at least part of the object, the first region having one or more first leak areas; inserting, using the hardware processor, a synthetic object obtained from a synthetic object database into the first frame of the video content, the synthetic object having a shape stored in the synthetic object database; selecting, using the hardware processor, a second region at least partially surrounding the synthetic object, the second region having one or more second leak areas; providing, using the hardware processor, the first frame including the first region and the second region to a first user for creating a first corrected frame; receiving, using the hardware processor, the first corrected frame including one or more first corrections by the first user to the one or more first leak areas of the first region, and one or more second corrections by the first user to the one or more second leak areas of the second region; determining, using the hardware processor, a quality of the one or more second corrections to the one or more second leak areas of the second region based on the shape of the synthetic object; and accepting, using the hardware processor, the one or more first corrections to the one or more first leak areas of the first region depending on the quality of the one or more second corrections to the one or more second leak areas of the second region.