Granted patent
Methods and systems for action recognition using poselet keyframes
- Number
- 10223580
- Published
- 2019-03-05
- Filed
- 2013-09-12
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Raptis; Michail, Sigal; Leonid
- CPC
- G06F18/254; G06V40/10; G06F18/24155; G06V40/20; G06V40/25; G06V40/171; G06V20/47
- Verdict
- Medium Notable software
- Source
- Google Patents · FreePatentsOnline
The keeper's note
Action recognition using poselet keyframes.
Abstract
Methods and systems for video action recognition using poselet keyframes are disclosed. An action recognition model may be implemented to spatially and temporally model discriminative action components as a set of discriminative keyframes. One method of action recognition may include the operations of selecting a plurality of poselets that are components of an action, encoding each of a plurality of video frames as a summary of the detection confidence of each of the plurality of poselets for the video frame, and encoding correlations between poselets in the encoded video frames.
Background
TECHNICAL FIELD(1) The present disclosure relates generally to action recognition techniques, and more particularly, some embodiments relate to methods for modeling human actions as a sequence of discriminative keyframes.DESCRIPTION OF THE RELATED ART(2) Most research on video-based action recognition modeling focuses on computing features over long temporal trajectories with many frames (e.g. 20 or 100 frame segments). This approach to action modeling, although descriptive, can be computationally expensive. Moreover, this approach to modeling may be sensitive to changes in action duration and dropped frames. Accordingly, an action recognition model that can effectively represent actions as a few key frames depicting key states of the action is desirable.BRIEF SUMMARY OF THE DISCLOSURE(3) According to various embodiments of the disclosed methods and systems, actions are modeled as a sequence of discriminative key frames. In one embodiment, a computer is configured to select a plurality of poselets that are components of an action, encode each of a plurality of video frames recorded by a video recording device as a summary of the detection confidence of each of the plurality of poselets for the video frame, and encode correlations between poselets in the encoded video frames. Optimal video frames may be selected based on the encoded video frames and encoded correlations. This process determines whether an action occurs in the video frames.(4) In one embodiment, each of the plu