- Number
- 9648303
- Published
- 2017-05-09
- Filed
- 2015-12-15
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Resch; Benjamin, Lensch; Hendrik, Pollefeys; Marc, Wang; Oliver, Hornung; Alexander Sorkine
- CPC
- G06T7/246; H04N13/264; G06V20/10; G06T7/73
- Verdict
- Low Notable software
- Source
- Google Patents · FreePatentsOnline
The keeper's note
Reconstructs 3D scenes from high-resolution, high-frame-rate video by determining camera pose and 3D feature point coordinates throughout the footage.
Abstract
Scenes reconstruction may be performed using videos that capture the scenes at high resolution and frame rate. Scene reconstruction may beassociated with determining camera orientation and/or location (“camera pose”) throughoutthe video, three-dimensional coordinates of feature points detected in frames of the video, and/or other information. Individual videos may have multiple frames. Feature points may be detected in, and tracked over, the frames. Estimations of camera pose may be made forindividual subsets of frames. One or more estimations of camera pose may be determined asfixed estimations. The estimated camera poses for the frames included in the subsets of frames may be updated based on the fixed estimations. Camera pose for frames not included in the subsets of frames may be determined to provide globally consistent camera poses and three-dimensional coordinates for feature points of the video.
Background
BRIEF DESCRIPTION OF THE DRAWINGS(1) FIG. 1 illustrates a system configured for facilitating three-dimensional reconstruction of scenes from videos, in accordance with one or more implementations.(2) FIG. 2 illustrates an exemplary implementation of a server employed in the system of FIG. 1.(3) FIG. 3 illustrates a graphical representation of a process of performing piecewise camera pose estimations over select subsets of frames of a given video, in accordance with one or more implementations.(4) FIG. 4 illustrates a method of facilitating three-dimensional reconstruction of scenes from videos, in accordance with one or more implementations.DETAILED DESCRIPTION(5) FIG. 1 illustrates a system 100 configured for facilitating three-dimensional reconstruction of scenes from videos, in accordance with one or more implementations. A video may comprise a recorded video, a live feed, and/or other audiovisual asset. A given video may have multiple frames, a sound track, and/or other components.(6) Insome implementations, the system 100 may comprise a server 102, one or more computing platforms 122, and/or other components. The server 102 may include one or more physical processors 104 configured by machine-readable instructions 106. Executing the machine-readable instructions 106 may cause the one or more physical processors 104 to facilitate three-dimensional reconstruction of scenes depicted in videos. The machine-readable instructions 106 may include one or more of a video compone
Claims
1. A system configured for facilitating three-dimensional reconstruction of scenes from videos, the systemcomprising: one or more physical processors configured by machine-readable instructions to: obtain a video having multiple frames, the video depicting a first scene captured from a perspective of a first camera; detect feature points of the first scene in individual frames of the video, a first set of feature points being detected in a first frame of the video; track features points over consecutive frames by correlating features points between consecutive frames, such that the first set of feature points is tracked from the first frame to one or more other frames by correlating the first set of feature points detected inthe first frame with other detections of the first set of feature points in the one or more other frames; make estimations of orientations and/or locations of the first camera in the first scene for individual frames within a first subset of frames of the video and a second subset of frames of the video, the second subset of frames comprising at least one frame that is not included in the first subset of frames, the estimations being based on the detected and tracked feature points of the first subset of frames and second subsets of frames, such that an estimation of a first orientation and/or location of the first camerais made for a second frame in the first subset of frames based on the detected and tracked feature points in the first subset of frames, and an estimation of a second orientation and/or location of the first camera is made for a third frame in the second subset of frames based on detected feature points in the second subset of frames; determine estimations of camera orientation/location that provide fixed estimations of orientation and/or location; and determine orientation and/or location of the first camera in the frames of the video based on the estimated first camera orientation and/or location, estimated second camera orientation and/or location, and the fixed estimations of orientation and/or location, the determined orientation and/or location of the first camera for the frames of the video facilitating three-dimensional reconstruction of the first scene of the video.
11. A method of facilitating three-dimensional reconstruction of scenes from videos, the method beingimplemented in a computer system comprises one or more physical processors and storage media storing machine-readable instructions, the method comprising: obtaining a video havingmultiple frames, the video depicting a first scene captured from the perspective of a first camera; detecting feature points of the first scene in individual frames of the video, including detecting a first set of feature points in a first frame of the video; tracking features points over consecutive frames by correlating features points between consecutiveframes, including tracking the first set of feature points from the first frame to one ormore other frames by correlating the first set of feature points detected in the first frame with other detections of the first set of feature points in the one or more other frames; making estimations of orientations and/or locations of the first camera in the first scene for individual frames within a first subset of frames of the video and a second subset of frames of the video, the second subset of frames comprising at least one frame that is not included in the first subset of frames, the estimations being based on the detected and tracked feature points of the first subset of frames and second subsets of frames, including making an estimation of a first orientation and/or location of the first camera fora second frame in the first subset of frames based on the detected and tracked feature points in the first subset of frames, and making an estimation of a second orientation and/or location of the first camera for a third frame in the second subset of frames based on detected feature points in the second subset of frames; determining estimations of camera orientation/location that provide fixed estimations of orientation and/or location; and determining orientation and/or location of the first camera in the frames of the video based on the estimated first camera orientation and/or location, estimated second camera orientation and/or location, and the fixed estimations of orientation and/or location, the orientation and/or location of the first camera for the frames of the video facilitating three-dimensional reconstruction of the first scene of the video.