Abstract: Methods are described for rendering a virtual object at a designated position in an input digital image corresponding to a perspective of a scene. In an embodiment, the method includes: estimating a set of lighting parameters using a lighting neural network; estimating a scene layout using a layout neural network; generating an environment texture map using a texture neural network using an input including the input digital image, the lighting parameters, and the scene layout; rendering the virtual object in a virtual scene constructed using the estimated lighting parameters, the scene layout, and the environment texture map; and compositing the rendered virtual object on the input digital image at the designated position. Corresponding systems and non-transitory computer-readable media are also described.