- Number
- 20220103764
- Published
- 2022-03-31
- Filed
- 2020-09-25
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Walsh; Peter, Devassy; Jayadas
- CPC
- G06T7/70; H04N5/272; G06T7/77
- Verdict
- Low Notable software
- Source
- Google Patents · FreePatentsOnline
The keeper's note
Model-based camera-tracking/occlusion-removal technique (VFX/AR).
Abstract
A system and method for model-based camera tracking and image occlusion removal for a camera viewing a sports field (or other scene) includes receiving a synthesized data set comprising at least one empty field image of the field, the empty field image with at least one occlusion graphic, and camera parameters corresponding to the empty field image, training a neural network model to estimate the empty field image and the corresponding camera parameters by providing the model with an input training image comprising the empty field image with occlusion graphic, and providing the model with model output targets comprising the empty field image and the corresponding camera parameters as targets for the model, receiving by the neural network model, alive input image comprising a view of the field with live occlusions, and providing by the neural network model, using trained model parameters, estimated live camera parameters or an estimated empty field image associated with the live input image.
Background
BACKGROUND
Many activities in broadcast and video production operations involve graphic insertion into moving video, each of which requires some form of camera tracking. Applications include broadcast enhancements for sports and other video productions. Types of graphic insertions include: live insertion; replay/post-production insertions; and, more recently, augmented reality insertions. All of these graphic insertions require an accurate model of the image formation process which can then be used with the generation of 3D graphics for insertion into the moving video. A spatially and temporally accurate model of the image formation process is necessary in order to match the insertion graphics to an actual scene with the required fidelity.
Previous solutions have included the use of: (i) electronic instrumentation on camera heads, lenses, and jibs; (ii) computer vision based video analysis, which utilizes explicit searches for known visual features; and (iii) video analysis in the context of augmented reality, which typically uses video analysis to find natural or artificial landmarks together with the use of inertial and magnetic sensors.
However, each of these camera tracking techniques have shortcomings. In particular, the instrumented camera approach requires a detailed calibration procedure to estimate the non-instrumented parameters, and requires on-site hardware set-up, support, and calibration requirements, and is very sensitive to vibration. The
Claims
1. A method for providing robust model-based camera tracking for a camera viewing a sports field, comprising: receiving a synthesized data set comprising at least one synthesized empty field image of the sports field, at least one of the synthesized empty field image with at least one synthesized occlusion graphic blocking at least a portion of the field, and synthesized camera parameters corresponding to camera parameters used to generate the synthesized empty field image; training a neural network model to estimate the corresponding synthesized camera parameters by providing the model with an input training image comprising the synthesized empty field image with synthesized occlusion graphic, and providing the model with model output targets comprising the synthesized empty field image and the corresponding synthesized camera parameters as targets for the model, and, when training is complete, the model providing trained model parameters; receiving by the neural network model, alive input image comprising a view of the field with live occlusions blocking portions of the field; and providing by the neural network model, using the trained model parameters, estimated live camera parameters corresponding to camera parameters used to generate the live input image. ||
12. A method for providing robust model-based camera tracking and occlusion removal for a camera viewing a sports field, comprising: receiving a synthesized data set comprising at least one synthesized empty field image of the field, at least one of the synthesized empty field image with at least one synthesized occlusion graphic blocking at least a portion of the field, and synthesized camera parameters corresponding to camera parameters used to generate the synthesized empty field image; training a neural network model to estimate the synthesized empty field image and the corresponding synthesized camera parameters by providing the model with an input training image comprising the synthesized empty field image with synthesized occlusion graphic, and providing the model with model output targets comprising the synthesized empty field image and the corresponding synthesized camera parameters as targets for the model, and, when training is complete, the model providing trained model parameters; receiving by the neural network model, alive input image comprising a view of the field with live occlusions blocking portions of the field; and providing by the neural network model, using the trained model parameters, estimated live camera parameters corresponding to camera parameters used to generate the live input image and an estimated empty field image associated with the live input image without any occlusions. ||
21. A method for providing occlusion removal for a camera viewing a scene, comprising: receiving a synthesized data set comprising at least one synthesized empty scene image of the scene, at least one of the synthesized empty scene image with at least one synthesized occlusion graphic blocking at least a portion of the scene, and synthesized camera parameters corresponding to camera parameters used to generate the synthesized empty scene image; training a neural network model to estimate the synthesized empty scene image and the corresponding synthesized camera parameters by providing the model with an input training image comprising the synthesized empty scene image with synthesized occlusion graphic, and providing the model with model output targets comprising the synthesized empty scene image and the corresponding synthesized camera parameters as targets for the model, and, when training is complete, the model providing trained model parameters; receiving by the neural network model, a live input image comprising a view of the scene with live occlusions blocking portions of the scene; and providing by the neural network model, using the trained model parameters, an estimated empty scene image associated with the live input image without any occlusions.