- Number
- 10726581
- Published
- 2020-07-28
- Filed
- 2015-06-18
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Wang; Oliver, Magnor; Marcus, Klose; Felix, Bazin; Jean-Charles, Sorkine Hornung; Alexander
- CPC
- H04N13/15; G06T7/90; H04N13/111
- Verdict
- Low Notable software
- Source
- Google Patents · FreePatentsOnline
The keeper's note
Scene-space video/3D processing rendering technique.
Abstract
There is provided a video processing system for use with a video having frames including a first frame and neighboring frames of the first frame. The system includes a memory storing a video processing application, and a processor. The processor is configured to execute the video processing application to sample scene points corresponding to an output pixel of the first frame of the frames of the video, the scene points including alternate observations of a same scene point from the neighboring frames of the first frame of the video, and filter the scene points corresponding to the output pixel to determine a color of the output pixel by calculating a weighted combination of the scene points corresponding to the output pixel.
Background
BACKGROUND(1) Many compelling video processing effects can be achieved if per pixel depth information and three-dimensional (3D) camera calibrations are known. Scene-space video processing, where pixels are processed according to their 3D positions, has many advantages over traditional image-space processing. For example, handling camera motion, occlusions, and temporal continuity entirely in two-dimensional (2D) image-space can in general be very challenging, while dealing with these issues in scene-space is simple. As scene-space information becomes more and more widely available due to advances in tools and mass market hardware devices, techniques that leverage depth information will play an important role in future video processing approaches. However, the success of such methods is highly dependent on the accuracy of the scene-space information.SUMMARY(2) The present disclosure is directed to systems and methods for scene-space video processing, substantially as shown in and/or described in connection with at least one of the figures, as set forth more completely in the claims.
Claims
1. A video processing system for use with a video having a plurality of frames including a first frame and a plurality of neighboring frames of the first frame, the system including: a display; a memory storing a video processing application; and a processor configured to execute the video processing application to: project a plurality of scene points of a selected number of the plurality of frames of the video produced by a camera into a three-dimensional (3D) scene-space using camera calibration parameters of the camera and depth values of the plurality of scene points, the selected number of the plurality of frames being produced by the camera and include the first frame and the plurality neighboring frames of the first frame, wherein projecting the plurality of scene points into the 3D scene-space creates a point cloud including a plurality of cloud points, wherein each cloud point of the plurality of cloud points corresponds to a projection of a scene point of the plurality of scene points, and wherein each of the plurality of scene points is a portion of a frame of video that is visible in a pixel of the frame of video when displayed on the display; sample, for each of a plurality of output pixels of the first frame, the projected plurality of scene points corresponding to an output pixel among the plurality of output pixels of the first frame, the projected plurality of scene points including alternate observations of a same scene point from the plurality of neighboring frames of the first frame; and filter the sampled plurality of scene points corresponding to the output pixel to determine a color of the output pixel by calculating a weighted combination of the sampled plurality of scene points to determine a video processing effect including one of a denoising, an object semi-transparency or a video inpainting, wherein filtering includes: identifying one or more erroneous observation points in the sampled plurality of scene points corresponding to the output pixel, wherein each of the one or more erroneous observation points corresponds to one of a scene point occlusion, incorrect three-dimensional (3D) information, or a moving object; and calculating the color of the output pixel by applying a weighting function to the sampled plurality of scene points, wherein the weighting function emphasizes scene points of the sampled plurality of scene points that are not the one or more erroneous observation points.
6. A method of video processing for use by a video processing system including a display, a memory, and a processor, the method comprising: projecting a plurality of scene points of a selected number of a plurality of frames of the video produced by a camera into a three-dimensional (3D) scene-space using camera calibration parameters of the camera and depth values of the plurality of scene points, the selected number of the plurality of frames being produced by the camera and include a first frame and a plurality of neighboring frames of the first frame, wherein projecting the plurality of scene points into the 3D scene-space creates a point cloud including a plurality of cloud points, wherein each cloud point of the plurality of cloud points corresponds to a projection of a scene point of the plurality of scene points, and wherein each of the plurality of scenes point is a portion of a frame of video that is visible in a pixel of the frame of video when displayed on the display; sampling, for each of a plurality of output pixels of the first frame, using the processor, the projected plurality of scene points corresponding to an output pixel among the plurality of output pixels of the first frame, the projected plurality of scene points including alternate observations of a same scene point from the plurality of neighboring frames of the first frame; and filtering, using the processor, the sampled plurality of scene points corresponding to the output pixel to determine a color of the output pixel by calculating a weighted combination of the sampled plurality of scene points to determine a video processing effect including one of a denoising, an object semi-transparency or a video inpainting, wherein filtering includes: identifying one or more erroneous observation points in the sampled plurality of scene points corresponding to the output pixel, wherein each of the one or more erroneous observation points corresponds to one of a scene point occlusion, incorrect three-dimensional (3D) information, or a moving object; and calculating the color of the output pixel by applying a weighting function to the sampled plurality of scene points, wherein the weighting function emphasizes scene points of the sampled plurality of scene points that are not the one or more erroneous observation points.