Outer Rim Archives
Archives · 2018 · 10102630

Granted patent

Video object tagging using segmentation hierarchy

Number
10102630
Published
2018-10-16
Filed
2015-04-21
Assignee
Disney Enterprises, Inc.
Inventors
Smolic; Aljoscha et al.
CPC
G06T7/12; G06F16/7867; G06T7/215; G11B27/00
Verdict
Set aside video tagging/content indexing
Source
Google Patents · FreePatentsOnline

Abstract

A system is provided for tagging an object in a video having a plurality of frames. The system includes a memory storing a segmentation hierarchy of a first frame of the plurality of frames and having a plurality of elements, a display, and a processor configured to display the first frame including the plurality of elements on the display, receive a first input selecting a first element of the plurality of elements displayed on the display, select a first region of the first frame based on the first input, display the first region of the first frame on the display, receive a second input from the user altering the first region of the first frame displayed on the display, and alter the first region by selecting a second region of the first frame based on the second input from the user and the segmentation hierarchy.

Background

BACKGROUND(1) Image and video segmentation is one of the most fundamental yet challenging problems in computer vision. Dividing an image into meaningful regions requires a high level interpretation of the image that cannot be satisfactorily solved by only looking for homogeneous areas in an image. In the era of big data and vast computing power, one approach to model high level interpretation of images has been to use powerful machine-learning tools on huge annotated databases. While significant advances have been made in recent years, automatic image segmentation is still far from providing accurate results in a generic scenario. The creator of a video may desire to add information or a link to an object in a video, and may wish the added information or link to remain associated with that object throughout a video sequence.SUMMARY(2) The present disclosure is directed to tagging objects in a video, substantially as shown in and/or described in connection with at least one of the figures, as set forth more completely in the claims.

Claims

1. A system comprising: a memory storing a segmentation hierarchy of each of a plurality of frames of a video, wherein the segmentation hierarchy is a segmentation of the plurality of frames of the video into a plurality of regions, each of the plurality of regions corresponding to an element of each of the plurality of frames, the plurality of frames including a first frame having a first plurality of regions corresponding to a first plurality of elements; a display; and a processor configured to: display the first frame including the first plurality of elements on the display; receive a first input from a user selecting a first element of the first plurality of elements displayed on the display; select a first region of the first plurality of regions of the first frame based on the first input from the user selecting the first element; display the first region of the first frame on the display; receive a second input from the user altering the first region of the first frame displayed on the display; alter the first region by selecting a second region of the first plurality of regions of the first frame based on the second input from the user and the segmentation hierarchy, wherein the second region comprises one or more masks; and propagate the one or more masks to a corresponding region of one or more other frames of the plurality of frames by: computing an optical flow linking pixels from the first frame to estimated positions to which the pixels have moved in the one or more other frames; and refining the one or more masks in the one or more other frames by adapting the selected second region according to each corresponding one of the segmentation hierarchies of the one or more other frames. 8. A method for use by a system having a display, a processor and a memory storing a segmentation hierarchy of each of a plurality of frames of a video, the segmentation hierarchy being a segmentation of the plurality of frames of the video into a plurality of regions, each of the plurality of regions corresponding to an element of each of the plurality of frames, the plurality of frames including a first frame having a first plurality of regions corresponding to a first plurality of elements, the method comprising: displaying, using the processor, the first frame including the first plurality of elements on the display; receiving, using the processor, a first input from a user selecting a first element of the first plurality of elements displayed on the display; selecting, using the processor, a first region of the first plurality of regions of the frame based on the first input from the user selecting the first element; displaying, using the processor, the first region of the first frame on the display; receiving, using the processor, a second input from the user altering the first region of the first frame displayed on the display; altering, using the processor, the first region by selecting a second region of the first plurality of regions of the first frame based on the second input from the user and the segmentation hierarchy, wherein the second region comprises one or more masks; and propagating, using the processor, the one or more masks to a corresponding region of one or more other frames of the plurality of frames by: computing an optical flow linking pixels from the first frame to estimated positions to which the pixels have moved in the one or more other frames; and refining the one or more masks in the one or more other frames by adapting the selected second region according to each corresponding one of the segmentation hierarchies of the one or more other frames.