Outer Rim Archives
Archives · 2022 · 11509818

Granted patent

Intelligent photography with machine learning

Number
11509818
Published
2022-11-22
Filed
2020-01-15
Assignee
Disney Enterprises, Inc.
Inventors
Adams; Mathew G., Kirkley; Jere M., Jenkins; Jason B.
CPC
G06N3/08; G06N3/09; H04N23/64; H04N23/611; H04N23/661; G06N3/0464; H04N23/90
Verdict
Set aside generic consumer photography ML feature
Source
Google Patents · FreePatentsOnline

Abstract

Embodiments of the present disclosure relate to intelligent photography with machine learning. Embodiments include receiving a video stream from a control camera. Embodiments include providing inputs to a trained machine learning model based on the video stream. Embodiments include determining, based on data output by the trained machine learning model in response to the inputs, at least a first time for capturing a first picture during a session. Embodiments include programmatically instructing a first camera to capture the first picture at the first time during the session.

Background

BACKGROUND Field of the Invention (1) The present invention generally relates to intelligent photography, and more specifically to techniques for using machine learning to determine indications for programmatically capturing photographs. Description of the Related Art (2) Photography generally involves analyzing a scene in order to determine optimal times for taking pictures, as well as identifying other optimal parameters, such as angles, exposures, and the like, for taking the pictures. When photographing live subjects, photographers generally observe the behavior of the subjects in order to determine photo-worthy events and optimal moments for capturing pictures. (3) In many settings, such as public places, photo-worthy events may occur frequently. For example, at theme parks, landmarks, sporting events, concerts and other public attractions, visitors are often photographed by both human and automated photographers and then provided with pictures of their visit, and may be given opportunities to purchase the pictures. Existing automated techniques, such as automatically capturing pictures at regular intervals for subsequent best-shot selection or post-processing to crop and improve photographs, may be inefficient and unresponsive to real-time circumstances. Furthermore, behavior of subjects may be unpredictable, making it more difficult to use systems based on fixed cues (e.g., regular intervals) for capturing pictures, as it is unlikely that such systems would successfull

Claims

1. A method performed by a first computing device, comprising: receiving a video stream from a control camera; providing inputs to a trained machine learning model based on the video stream; determining, based on data output by the trained machine learning model in response to the inputs, at least a first time for capturing a first picture during a session; identifying, based on the data output by the trained machine learning model in response to the inputs, a first camera from a plurality of cameras for capturing the first picture; and programmatically instructing the first camera to capture the first picture at the first time during the session. || 8. A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform a method, the method comprising: receiving a video stream from a control camera; providing inputs to a trained machine learning model based on the video stream; determining, based on data output by the trained machine learning model in response to the inputs, at least a first time for capturing a first picture during a session; identifying, based on the data output by the trained machine learning model in response to the inputs, a first camera from a plurality of cameras for capturing the first picture; and programmatically instructing the first camera to capture the first picture at the first time during the session. || 16. A system, comprising: one or more processors; and a non-transitory computer-readable medium comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform a method, the method comprising: receiving a video stream from a control camera; providing inputs to a trained machine learning model based on the video stream; determining, based on data output by the trained machine learning model in response to the inputs, at least a first time for capturing a first picture during a session; identifying, based on the data output by the trained machine learning model in response to the inputs, a first camera from a plurality of cameras for capturing the first picture; and programmatically instructing the first camera to capture the first picture at the first time during the session. || 20. A method performed by a first computing device, comprising: receiving a video stream from a control camera; providing inputs to a trained machine learning model based on the video stream, wherein the inputs to the trained machine learning model are based on one or more of: a body position of a subject present in the video stream; a facial expression of the subject present in the video stream; or a gaze direction of the subject present in the video stream; determining, based on data output by the trained machine learning model in response to the inputs, at least a first time for capturing a first picture during a session; and programmatically instructing a first camera to capture the first picture at the first time during the session.