Outer Rim Archives
Archives · 2018 · 20180262660

Application (pre-grant publication)

METHOD AND SYSTEM FOR MIMICKING HUMAN CAMERA OPERATION

Number
20180262660
Published
2018-09-13
Filed
2018-05-09
Assignee
Disney Enterprises, Inc.
Inventors
CARR; George Peter; CHEN; Jianhui
CPC
H04N17/002; H04N5/2228
Verdict
Medium Hardware
Source
Google Patents · FreePatentsOnline

The keeper's note

Robotic camera mimicking human operation (PGPUB dup).

Abstract

The disclosure provides an approach for mimicking human camera operation with an autonomous camera system. In one embodiment, camera planning is formulated as a supervised regression problem in which an automatic broadcasting application receives one video input captured by a human-operated camera and another video input captured by a stationary camera with a wider field of view. The automatic broadcasting application extracts feature vectors and pan-tilt-zoom states from the stationary camera and the human-operated camera, respectively, and learns a regressor which takes as input such feature vectors and outputs pan-tilt-zoom settings predictive of what the human camera operator would choose. The automatic broadcasting application may then apply the learned regressor on newly captured video to obtain planned pan-tilt-zoom settings and control an autonomous camera to achieve the planned settings to record videos which resemble the work of a human operator in similar situations.

Background

BACKGROUNDField

This disclosure provides techniques for automatically capturing video. More specifically, embodiments of this disclosure present techniques for mimicking human camera operators in capturing a video.Description of the Related Art

Automatic broadcasting, in which autonomous camera systems capture video, can make small events, such as lectures and amateur sporting competitions, available to much larger audiences. Autonomous camera systems generally need the capability to sense the environment, decide where to point a camera (or cameras) when recording, and ensure the cameras remain fixated on intended targets. Traditionally, autonomous camera systems follow an object-tracking paradigm, such as “follow the lecturer,” and implement camera planning (i.e., determining where the camera should look) by smoothing the data from the object tracking, which tends to be noisy. Such autonomous camera systems typically include hand-coded equations which determine where to point each camera. One problem with such systems is that, unlike human camera operators, hand-coded autonomous camera systems cannot anticipate action and frame their shots with sufficient “lead room.” As a result, the output videos produced by such systems tend to look robotic, particularly for dynamic activities such as sporting events.SUMMARY

One embodiment of this disclosure provides a computer implemented method for building a model to control a first device. The method generally includ

Claims

1. A method for building a model to control a first device, comprising: receiving, as input: demonstration data corresponding to human operation of a second device used to perform a demonstration, wherein the second device is a camera and wherein the demonstration data includes a first video captured by the camera under control of the human operator; and environmental sensory data associated with the demonstration data; determining device settings of the second device from the demonstration data; extracting, from the environmental sensory data, feature vectors describing at least locations of objects in the environment; training, based on the determined device settings and the extracted feature vectors, a regressor which takes additional feature vectors as inputs, and outputs planned device settings for operating the first device; and instructing the first device to attain the planned device settings output by the trained regressor. 11. A non-transitory computer-readable storage medium storing a program, which, when executed by a processor performs operations for building a model to control a first device, the operations comprising: receiving, as input: demonstration data corresponding to human operation of a second device used to perform a demonstration, wherein the second device is a camera and wherein the demonstration data includes a first video captured by the camera under control of the human operator; and environmental sensory data associated with the demonstration data; determining device settings of the second device from the demonstration data; extracting, from the environmental sensory data, feature vectors describing at least locations of objects in the environment; training, based on the determined device settings and the extracted feature vectors, a regressor which takes additional feature vectors as inputs, and outputs planned device settings for operating the first device; and instructing the first device to attain the planned device settings output by the trained regressor. 20. A system, comprising: a first data capture device; a second data capture device; a processor; and a memory, wherein the memory includes an application program configured to perform operations for building a model to control the first data capture device, the operations comprising: receiving, as input: demonstration data corresponding to human operation of the second data capture device used to perform a demonstration, wherein the second device is a camera and wherein the demonstration data includes a first video captured by the camera under control of the human operator; and environmental sensory data associated with the demonstration data; determining device settings of the second data capture device from the demonstration data; extracting, from the environmental sensory data, feature vectors describing at least locations of objects in the environment; training, based on the determined device settings and the extracted feature vectors, a regressor which takes additional feature vectors as inputs, and outputs planned device settings for operating the first device; and instructing the first data capture device to attain the planned device settings output by the trained regressor.