- Number
- 10200618
- Published
- 2019-02-05
- Filed
- 2015-10-12
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Carr; George Peter, Chen; Jianhui, Yue; Yisong
- CPC
- G06T7/277; H04N23/69
- Verdict
- Low Notable software
- Source
- Google Patents · FreePatentsOnline
The keeper's note
Automatic device operation/object tracking via smooth predictors.
Abstract
The disclosure provides an approach for predicting trajectories for real-time capture of video and object tracking, while adhering to smoothness constraints so that predictions are not excessively jittery. In one embodiment, a temporally consistent search and learn (TC-SEARN) algorithm is applied to train a regressor for camera planning. A automatic broadcasting application first receives video input captured by a human-operated camera and another video input captured by a stationary camera with a wide field of view. The automatic broadcasting application extracts feature vectors and pan-tilt-zoom states from the stationary camera input and human-operated camera input, respectively. The automatic broadcasting application further applies the TC-SEARN algorithm to learn a sequential regressor for predicting camera trajectories, based on the extracted feature vectors and pan-tilt-zoom states. The TC-SEARN algorithm itself is able to learn the regressor using a loss function which enables decision trees to reason about spatiotemporal smoothness via an autoregressive function.
Background
BACKGROUND(1) Field of the Invention(2) This disclosure provides techniques for automatically operating a device or tracking a device's or an object's states. More specifically, embodiments of this disclosure present techniques for learning smooth predictors for automatic device operation and object tracking.(3) Description of the Related Art(4) Automatic broadcasting, in which autonomous camera systems capture video, can make small events, such as lectures and amateur sporting competitions, available to much larger audiences. Autonomous camera systems generally need the capability to sense the environment, decide where to point a camera (or cameras) when recording, and ensure the cameras remain fixated on intended targets. Traditionally, autonomous camera systems follow an object-tracking paradigm, such as “follow the lecturer,” and implement camera planning (i.e., determining where the camera should look) by smoothing the data from the object tracking, which tends to be noisy. Such autonomous camera systems typically include hand-coded equations which determine where to point each camera. One problem with such systems is that, unlike human camera operators, hand-coded autonomous camera systems cannot anticipate action and frame their shots with sufficient “lead room.” As a result, the output videos produced by such systems tend to look robotic, particularly for dynamic activities such as sporting events. Another problem is that automatically generated camera plans can be ji
Claims
1. A computer-implemented method for operating a first device, comprising: receiving operational data corresponding to manual operation of a second device at a first operating site; receiving sensory data captured using at least one device at the first operating site, the operational data and the sensory data being associated with a same temporal duration; extracting feature vectors from the sensory data; training a regressor based, at least in part, on the operational data and the feature vectors, thereby generating a trained regressor, wherein the training employs a temporally consistent search and learn (TC-SEARN) method and includes minimizing a spatiotemporal loss function; and automatically controlling operation of the first device at a second operating site using the trained regressor, into which additional feature vectors extracted from additional sensory data captured using at least one device at the second operating site are input, wherein the operational data includes settings of the second device, and wherein each of the sensory data and the additional sensory data include at least one of video data, radio frequency identification (RFID) tracking data, radar data, and scoreboard data.
9. A computer-implemented method for tracking a first object, comprising: receiving data representing states of a second object, wherein the states of the second object include locations of the second object; receiving sensory data that is captured using at least one device and associated with a same temporal duration as the states of the second object, wherein the sensory data includes at least one of video data, radio frequency identification (RFID) tracking data, radar data, and scoreboard data; extracting feature vectors from the sensory data; training a regressor based, at least in part, on the second object state data and the feature vectors, thereby generating a trained regressor, wherein the training employs a temporally consistent search and learn (TC-SEARN) method and minimizes a spatiotemporal loss function; and automatically tracking the first object using the trained regressor, into which additional feature vectors extracted from additional sensory data captured using at least one device are input.
16. A non-transitory computer-readable storage medium storing a program, which, when executed by a processor performs operations for operating a first device, the operations comprising: receiving operational data corresponding to manual operation of a second device at a first operating site. receiving sensory data captured using at least one device at the first operating site, the operational data and the sensory data being associated with a same temporal duration; extracting feature vectors from the sensory data; training a regressor based, at least in part, on the operational data and the feature vectors, thereby generating a trained regressor, wherein the training employs a temporally consistent search and learn (TC-SEARN) method and includes minimizing a spatiotemporal loss function; and automatically controlling operation of the first device at a second operating site using the trained regressor, into which additional feature vectors extracted from additional sensory data captured using at least one device at the second operating site are input, wherein the operational data includes settings of the second device, and wherein each of the sensory data and the additional sensory data include at least one of video data, radio frequency identification (RFID) tracking data, radar data, and scoreboard data.