Outer Rim Archives
Archives · 2018 · 10057561

Granted patent

Systems and methods for facilitating three-dimensional reconstruction of scenes from videos

Number
10057561
Published
2018-08-21
Filed
2017-05-01
Assignee
Disney Enterprises, Inc.; ETH Zurich
Inventors
Resch; Benjamin; Lensch; Hendrik; Pollefeys; Marc; Wang; Oliver
CPC
H04N13/264; G06V20/10; G06T7/73; G06T7/246
Verdict
Medium Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

3D scene reconstruction from video.

Abstract

Scenes reconstruction may be performed using videos that capture the scenes at high resolution and frame rate. Scene reconstruction may be associated with determining camera orientation and/or location (“camera pose”) throughout the video, three-dimensional coordinates of feature points detected in frames of the video, and/or other information. Individual videos may have multiple frames. Feature points may be detected in, and tracked over, the frames. Estimations of camera pose may be made for individual subsets of frames. One or more estimations of camera pose may be determined as fixed estimations. The estimated camera poses for the frames included in the subsets of frames may be updated based on the fixed estimations. Camera pose for frames not included in the subsets of frames may be determined to provide globally consistent camera poses and three-dimensional coordinates for feature points of the video.

Background

FIELD OF THE DISCLOSURE(1) This disclosure relates to three-dimensional reconstruction of scenes from videos.BACKGROUND(2) Structure from Motion (SfM), Simultaneous Localization and Mapping (SLAM), bundle adjustment, and/or other techniques may be used for three-dimensional scene reconstruction. Three-dimensional scene reconstruction may include determining one or more of a camera location, camera orientation, and/or scene geometry from images. One or more techniques may involve feature point detection within individual images and/or feature point tracked over multiple images. Feature point detection and/or tracking may be accomplished using one or more of Scale-Invariant Feature Transform (SIFT), Speeded Up Robust Features (SURF), Orientated Features From Accelerated Segment Test and Rotated Binary Robust Independent Elementary Features (ORB), Kanade-Lucas-Tomasi (KLT), and/or other techniques for detecting and/or tracking feature points in images. One or more feature point detection and/or tracking processes may return feature point descriptors and/or other information. By way of non-limiting example, a feature point descriptor may comprise information including one or more of a spatial histogram of the image gradients and/or other information, a sum of the Haar wavelet response around feature point, an intensity distribution of pixels within a region surrounding a feature point, and/or other information.(3) A bundle adjustment process may be used to determine one or more o

Claims

1. A system configured to facilitate three-dimensional reconstruction of scenes depicted in videos, the system comprising: one or more physical processors configured by machine-readable instructions to: obtain a video having multiple frames, the video depicting a first scene captured from a first camera, the first scene including feature points within individual frames of the video, the feature points being tracked over consecutive frames by correlating the feature points between the consecutive frames, a first frame of the video including a first set of feature points, the first set of feature points being tracked from the first frame to one or more other frames by correlating the first set of feature points within the first frame with the first set of feature points present within the one or more other frames; make estimations of orientation and/or location of the first camera in the first scene for individual frames within a first subset of frames of the video and a second subset of frames of the video, the second subset of frames comprising at least one frame not included in the first subset of frames, the estimations being based on the detected and tracked feature points of the first subset of frames and second subsets of frames, such that an estimation of a first orientation and/or location of the first camera is made for a second frame in the first subset of frames based on the detected and tracked feature points in the first subset of frames, and an estimation of a second orientation and/or location of the first camera is made for a third frame in the second subset of frames based on detected feature points in the second subset of frames; determine estimations of camera orientation and/or location which provide fixed estimations of orientation and/or location; and determine orientation and/or location of the first camera in the frames of the video based on the estimated first camera orientation and/or location, estimated second camera orientation and/or location, and the fixed estimations of orientation and/or location, the determined orientation and/or location of the first camera for the frames of the video facilitating three-dimensional reconstruction of the first scene of the video. 11. A method of facilitating three-dimensional reconstruction of scenes depicted in videos, the method being implemented in a computer system comprising one or more physical processors and storage media storing machine-readable instructions, the method comprising: obtaining a video having multiple frames, the video depicting a first scene captured from a first camera, the first scene including feature points within individual frames of the video, the feature points being tracked over consecutive frames by correlating the feature points between the consecutive frames, a first frame of the video including a first set of feature points, the first set of feature points being tracked from the first frame to one or more other frames by correlating the first set of feature points within the first frame with the first set of feature points present within the one or more other frames; making estimations of orientation and/or location of the first camera in the first scene for individual frames within a first subset of frames of the video and a second subset of frames of the video, the second subset of frames comprising at least one frame not included in the first subset of frames, the estimations being based on the detected and tracked feature points of the first subset of frames and second subsets of frames, including making an estimation of a first orientation and/or location of the first camera for a second frame in the first subset of frames based on the detected and tracked feature points in the first subset of frames, and making an estimation of a second orientation and/or location of the first camera for a third frame in the second subset of frames based on detected feature points in the second subset of frames; determining estimations of camera orientation/location which provide fixed estimations of orientation and/or location; and determining orientation and/or location of the first camera in the frames of the video based on the estimated first camera orientation and/or location, estimated second camera orientation and/or location, and the fixed estimations of orientation and/or location, the orientation and/or location of the first camera for the frames of the video facilitating three-dimensional reconstruction of the first scene of the video.