Abstract: Systems and methods for enhancing a composite video sequence created by integrating a real-world object into a synthetic environment. An initial composite video sequence, which may contain visual inconsistencies, is received. Object appearance information, derived from an original source recording of the object, is also received to guide the enhancement process. A generative model processes the initial composite video, conditioned on both the initial composite's content and the object appearance information. This process generates an enhanced video sequence wherein the visual integration of the object into the environment is improved. Enhancements include corrections to lighting, color, contrast, and shadows, as well as generative stabilization of noisy camera motion and correction of viewpoint discrepancies. Critically, the process preserves the identity of the real-world object.
Abstract: Systems and methods for integrating a sequence of 2D-images of an object into a synthetic 3D-scene. A sequence of 2D-images of a target object is captured using a smart device such as a smartphone. The object is extracted from the image sequence, and converted into a corresponding sequence of flat-surfaced 3D-renderable objects that are placed in a synthetic 3D-scene. Movement and orientation of the smart device are captured and translated into corresponding viewing points in the 3D-scene, in which the viewing points are then used to 3D render the 3D-scene, together with the flat-surfaced 3D-renderable objects now embedded therewith, into a video sequence showing the target object as an integral part of the 3D-scene. Other effects such as lighting, shadowing and reflections are rendered in conjunction with the flat-surfaced 3D-renderable objects so as to further enhance an illusion that the target object is an integral part of the 3D-scene.