Patents by Inventor Igor Vasiljevic

Igor Vasiljevic has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Patent number: 12670211
    Abstract: A method for determining a complexity of a natural language query includes converting a first natural language query into executable program code, the first natural language query being a query for a first video to be answered by one or more video question answering (VideoQA) models. The method also includes generating, via a complexity model, an abstract syntax tree (AST) based on the executable program code. The method further includes determining, via the complexity model, a complexity of the first natural language query based on quantity of subtrees, from a group of subtrees, that are present in the AST.
    Type: Grant
    Filed: January 28, 2025
    Date of Patent: June 30, 2026
    Assignees: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHA, THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIVERSITY
    Inventors: Cristobal Eyzaguirre, Igor Vasiljevic, Achal Dave, Jiajun Wu, Thomas Kollar, Juan Carlos Niebles, Pavel Tokmakov
  • Patent number: 12633038
    Abstract: System, methods, and other embodiments described herein relate to generating an image by interpolating features estimated from a learning model. In one embodiment, a method includes sampling three-dimensional (3D) points of a light ray that crosses a frustum space associated with a single-view camera, the 3D points reflecting depth estimates derived from data that the single-view camera generates for a scene. The method also includes deriving feature values for the 3D points using tri-linear interpolation across feature planes of the frustum space, the feature planes being estimated by a learning model. The method also includes inferring an image in two dimensions (2D) by translating the feature values and compositing the data with volumetric rendering for the scene. The method also includes executing a control task by a controller using the image.
    Type: Grant
    Filed: March 29, 2023
    Date of Patent: May 19, 2026
    Assignees: Toyota Research Institute, Inc., Toyota Jidosha Kabuhiki Kaisha, Toyota Technological Institute at Chicago
    Inventors: Jiading Fang, Vitor Guizilini, Igor Vasiljevic, Rares A. Ambrus, Gregory Shakhnarovich, Matthew R. Walter, Adrien David Gaidon
  • Publication number: 20260094449
    Abstract: Systems and methods described herein relate to self-supervised scale-aware learning of camera extrinsic parameters. One embodiment processes instantaneous velocity between a target image and a context image captured by a first camera; jointly training a depth network and pose network based on scaling by the instantaneous velocity; produce depth map using the depth network; produce ego-motion of the first camera using the pose network; generate synthesized image from the target image using a reprojection operation based on the depth map, the ego-motion, the context image and camera intrinsics; determine photometric loss by comparing the synthesized image to the target image; generate photometric consistency constraint using a gradient from the photometric loss; determine pose consistency constraint between the first camera and a second camera; and optimize the photometric consistency constraint, the pose consistency constraint, the depth network and the pose network to generate estimated extrinsic parameters.
    Type: Application
    Filed: December 7, 2025
    Publication date: April 2, 2026
    Applicants: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHA
    Inventors: TAKAYUKI KANAI, Vitor Campagnolo Guizilini, Rares A. Ambrus, Adrien Gaidon, Igor Vasiljevic
  • Patent number: 12530835
    Abstract: An example method includes generating embeddings of image data that includes multiple images, where each image has a different viewpoints of a scene, generating a latent space and a decoder, wherein the decoder receives embeddings as input to generate an output viewpoint, for each viewpoint in the image data, determining a volumetric rendering view synthesis loss and a multi-view photometric loss, and applying an optimization algorithm to the latent space and the decoder over a number of epochs until the volumetric rendering view synthesis loss is within a volumetric threshold and the multi-view photometric loss is within a multi-view threshold.
    Type: Grant
    Filed: August 3, 2023
    Date of Patent: January 20, 2026
    Assignees: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki Kaisha, Massachusetts Institute of Technology
    Inventors: Vitor Guizilini, Rares A. Ambrus, Jiading Fang, Sergey Zakharov, Vincent Sitzmann, Igor Vasiljevic, Adrien Gaidon
  • Patent number: 12524952
    Abstract: Systems and methods described herein support enhanced computer vision capabilities which may be applicable to, for example, autonomous vehicle operation. An example method includes generating a latent space and a decoder based on image data that includes multiple images, where each image has a different viewing frame of a scene. The method also includes generating a volumetric embedding that is representative of a novel viewing frame of the scene. The method includes decoding, with the decoder, the latent space using cross-attention with the volumetric embedding, and generating a novel viewing frame of the scene based on an output of the decoder.
    Type: Grant
    Filed: August 3, 2023
    Date of Patent: January 13, 2026
    Assignees: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki Kaisha, Massachusetts Institute of Technology
    Inventors: Vitor Guizilini, Rares A. Ambrus, Jiading Fang, Sergey Zakharov, Vincent Sitzmann, Igor Vasiljevic, Adrien Gaidon
  • Patent number: 12524894
    Abstract: A method for scale-aware depth estimation using multi-camera projection loss is described. The method includes determining a multi-camera photometric loss associated with a multi-camera rig of an ego vehicle. The method also includes training a scale-aware depth estimation model and an ego-motion estimation model according to the multi-camera photometric loss. The method further includes predicting a 360° point cloud of a scene surrounding the ego vehicle according to the scale-aware depth estimation model and the ego-motion estimation model. The method also includes planning a vehicle control action of the ego vehicle according to the 360° point cloud of the scene surrounding the ego vehicle.
    Type: Grant
    Filed: June 5, 2024
    Date of Patent: January 13, 2026
    Assignees: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA TECHNOLOGICAL INSTITUTE AT CHICAGO
    Inventors: Vitor Guizilini, Rares Andrei Ambrus, Adrien David Gaidon, Igor Vasiljevic, Gregory Shakhnarovich
  • Patent number: 12511910
    Abstract: Systems and methods described herein relate to self-supervised scale-aware learning of camera extrinsic parameters. One embodiment processes instantaneous velocity between a target image and a context image captured by a first camera; jointly training a depth network and pose network based on scaling by the instantaneous velocity; produce depth map using the depth network; produce ego-motion of the first camera using the pose network; generate synthesized image from the target image using a reprojection operation based on the depth map, the ego-motion, the context image and camera intrinsics; determine photometric loss by comparing the synthesized image to the target image; generate photometric consistency constraint using a gradient from the photometric loss; determine pose consistency constraint between the first camera and a second camera; and optimize the photometric consistency constraint, the pose consistency constraint, the depth network and the pose network to generate estimated extrinsic parameters.
    Type: Grant
    Filed: September 18, 2023
    Date of Patent: December 30, 2025
    Assignees: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHA
    Inventors: Takayuki Kanai, Vitor Campagnolo Guizilini, Rares A. Ambrus, Adrien Gaidon, Igor Vasiljevic
  • Publication number: 20250373770
    Abstract: Systems and methods for enhanced computer vision capabilities, particularly including depth synthesis, which may be applicable to autonomous vehicle operation are described. A vehicle may be equipped with a geometric scene representation (GSR) architecture for synthesizing depth views at arbitrary viewpoints. The GSR architecture synthesizes depth views enable advanced functions, including depth interpolation and depth extrapolation. The GSR architecture implements functions (i.e., depth interpolation, depth extrapolation) that are useful for various computer vision applications for autonomous vehicles, such as predicting depth maps from unseen locations. For example, a vehicle includes a processor device synthesizing depth views at multiple viewpoints, where the multiple viewpoints are from image data of a surrounding environment for the vehicle.
    Type: Application
    Filed: August 19, 2025
    Publication date: December 4, 2025
    Applicants: TOYOTA TECHNOLOGICAL INSTITUTE AT CHICAGO, TOYOTA JIDOSHA KABUSHIKI KAISHA, TOYOTA RESEARCH INSTITUTE, INC.
    Inventors: VITOR GUIZILINI, IGOR VASILJEVIC, Adrien D. GAIDON, Greg SHAKHNAROVICH, Matthew WALTER, Jiading FANG, Rares A. AMBRUS
  • Patent number: 12488483
    Abstract: A method of generating additional supervision data to improve learning of a geometrically-consistent latent scene representation with a geometric scene representation architecture is provided. The method includes receiving, with a computing device, a latent scene representation encoding a pointcloud from images of a scene captured by a plurality of cameras each with known intrinsics and poses, generating a virtual camera having a viewpoint different from viewpoints of the plurality of cameras, projecting information from the pointcloud onto the viewpoint of the virtual camera, and decoding the latent scene representation based on the virtual camera thereby generating an RGB image and depth map corresponding to the viewpoint of the virtual camera for implementation as additional supervision data.
    Type: Grant
    Filed: February 16, 2023
    Date of Patent: December 2, 2025
    Assignees: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki Kaisha, Toyota Technological Institute at Chicago
    Inventors: Vitor Guizilini, Igor Vasiljevic, Adrien D. Gaidon, Jiading Fang, Gregory Shakhnarovich, Matthew R. Walter, Rares A. Ambrus
  • Publication number: 20250355934
    Abstract: A method for determining a complexity of a natural language query includes converting a first natural language query into executable program code, the first natural language query being a query for a first video to be answered by one or more video question answering (VideoQA) models. The method also includes generating, via a complexity model, an abstract syntax tree (AST) based on the executable program code. The method further includes determining, via the complexity model, a complexity of the first natural language query based on quantity of subtrees, from a group of subtrees, that are present in the AST.
    Type: Application
    Filed: January 28, 2025
    Publication date: November 20, 2025
    Applicants: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHA, THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIVERSITY
    Inventors: Cristobal EYZAGUIRRE, Igor VASILJEVIC, Achal DAVE, Jiajun WU, Thomas KOLLAR, Juan Carlos NIEBLES, Pavel TOKMAKOV
  • Publication number: 20250307616
    Abstract: A method may include receiving parameters associated with a pre-trained transformer trained on first training data, modifying an architecture of the pre-trained transformer to generate a modified transformer, the modified transformer replacing a dot-product softmax attention layer with a linear kernel squared dot product attention layer utilizing Group Normalization, receiving second training data, and training the modified transformer based on the training data.
    Type: Application
    Filed: January 31, 2025
    Publication date: October 2, 2025
    Applicants: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki Kaisha
    Inventors: Jean Mercat, Igor Vasiljevic, Sedrick Keh, Achal Dave, Kushal Arora, Thomas Kollar
  • Publication number: 20250307633
    Abstract: A method may include receiving parameters associated with a pre-trained transformer trained on first training data, modifying an architecture of the pre-trained transformer to generate a modified transformer, the modified transformer replacing a dot-product softmax attention layer with a linear kernel dot product attention layer utilizing Group Normalization, receiving second training data, and training the modified transformer based on the training data.
    Type: Application
    Filed: January 31, 2025
    Publication date: October 2, 2025
    Applicants: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki Kaisha
    Inventors: Igor Vasiljevic, Jean Mercat, Sedrick Keh, Achal Dave, Kushal Arora, Thomas Kollar
  • Patent number: 12430840
    Abstract: Systems and methods for enhanced computer vision capabilities, particularly including depth synthesis, which may be applicable to autonomous vehicle operation are described. A vehicle may be equipped with a geometric scene representation (GSR) architecture for synthesizing depth views at arbitrary viewpoints. The GSR architecture synthesizes depth views enable advanced functions, including depth interpolation and depth extrapolation. The GSR architecture implements functions (i.e., depth interpolation, depth extrapolation) that are useful for various computer vision applications for autonomous vehicles, such as predicting depth maps from unseen locations. For example, a vehicle includes a processor device synthesizing depth views at multiple viewpoints, where the multiple viewpoints are from image data of a surrounding environment for the vehicle.
    Type: Grant
    Filed: January 19, 2023
    Date of Patent: September 30, 2025
    Assignees: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHA, TOYOTA TECHNOLOGICAL INSTITUTE AT CHICAGO
    Inventors: Vitor Guizilini, Igor Vasiljevic, Adrien D. Gaidon, Greg Shakhnarovich, Matthew Walter, Jiading Fang, Rares A. Ambrus
  • Publication number: 20250272863
    Abstract: Systems and methods for self-supervised learning for visual odometry are provided. An example method may comprise: (1) using a keypoint network to generate a keypoint matrix for a target image captured by a camera and a keypoint matrix for a context image captured by the camera, each keypoint matrix comprising keypoints of its respective image; (2) using a neural camera model to predict a pixel-wise ray surface for the target image, wherein the predicted pixel-wise ray surface associates a respective pixel in the target image with a corresponding direction; (3) using the generated keypoint matrices and the predicted pixel-wise ray surface to: (a) lift 2D keypoints of the target image to 3D keypoints, and (b) project the 3D keypoints of the target image into the context image; and (4) computing a geometric loss based on differences between the projected keypoints of the target image and the keypoints of the context image.
    Type: Application
    Filed: May 12, 2025
    Publication date: August 28, 2025
    Inventors: VITOR GUIZILINI, IGOR VASILJEVIC, RARES A. AMBRUS, SUDEEP PILLAI, ADRIEN GAIDON
  • Publication number: 20250265725
    Abstract: A method may include receiving a plurality of pairs of images of a scene captured by a camera, receiving a pose of the camera when each of the pairs of images of the scene were captured, determining depth values of the scene for each pair of images, and training a neural network to receive a pose of a camera with respect to the scene, and output a geometry of the scene and an appearance of the scene with respect to the pose. The plurality of pairs of images, the poses of the camera, and the depth values may be used as training data to train the neural network. The neural network may include a first component that receives the pose as input, and outputs the geometry of the scene and an embedding, and a second component that receives the embedding as input, and outputs the appearance of the scene.
    Type: Application
    Filed: January 31, 2025
    Publication date: August 21, 2025
    Applicants: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki Kaisha, Carnegie Mellon University
    Inventors: Igor Vasiljevic, Sergey Zakharov, Vitor Guizilini, Rares Ambrus, Arkadeep Narayan Chaudhury
  • Patent number: 12333750
    Abstract: Systems and methods for self-supervised learning for visual odometry using camera images, may include: estimating correspondences between keypoints of a target camera image and keypoints of a context camera image; based on the keypoint correspondences, lifting a set of 2D keypoints to 3D, using a neural camera model; and projecting the 3D keypoints into the context camera image using the neural camera model. Some embodiments may use the neural camera model to achieve the lifting and projecting of keypoints without a known or calibrated camera model.
    Type: Grant
    Filed: April 17, 2022
    Date of Patent: June 17, 2025
    Assignee: TOYOTA RESEARCH INSTITUTE, INC.
    Inventors: Vitor Guizilini, Igor Vasiljevic, Rares A. Ambrus, Sudeep Pillai, Adrien Gaidon
  • Patent number: 12293548
    Abstract: Systems, methods, and other embodiments described herein relate to estimating scaled depth maps by sampling variational representations of an image using a learning model. In one embodiment, a method includes encoding data embeddings by a learning model to form conditioned latent representations using attention networks, the data embeddings including features about an image from a camera and calibration information about the camera. The method also includes computing a probability distribution of the conditioned latent representations by factoring scale priors. The method also includes sampling the probability distribution to generate variations for the data embeddings. The method also includes estimating scaled depth maps of a scene from the variations at different coordinates using the attention networks.
    Type: Grant
    Filed: October 13, 2023
    Date of Patent: May 6, 2025
    Assignees: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki Kaisha
    Inventors: Vitor Campagnolo Guizilini, Igor Vasiljevic, Dian Chen, Adrien David Gaidon, Rares A. Ambrus
  • Publication number: 20250095380
    Abstract: Systems and methods described herein relate to self-supervised scale-aware learning of camera extrinsic parameters. One embodiment processes instantaneous velocity between a target image and a context image captured by a first camera; jointly training a depth network and pose network based on scaling by the instantaneous velocity; produce depth map using the depth network; produce ego-motion of the first camera using the pose network; generate synthesized image from the target image using a reprojection operation based on the depth map, the ego-motion, the context image and camera intrinsics; determine photometric loss by comparing the synthesized image to the target image; generate photometric consistency constraint using a gradient from the photometric loss; determine pose consistency constraint between the first camera and a second camera; and optimize the photometric consistency constraint, the pose consistency constraint, the depth network and the pose network to generate estimated extrinsic parameters.
    Type: Application
    Filed: September 18, 2023
    Publication date: March 20, 2025
    Applicants: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHA
    Inventors: TAKAYUKI KANAI, Vitor Campagnolo Guizilini, Rares A. Ambrus, Adrien Gaidon, Igor Vasiljevic
  • Patent number: 12175708
    Abstract: Systems and methods described herein relate to self-supervised learning of camera intrinsic parameters from a sequence of images. One embodiment produces a depth map from a current image frame captured by a camera; generates a point cloud from the depth map using a differentiable unprojection operation; produces a camera pose estimate from the current image frame and a context image frame; produces a warped point cloud based on the camera pose estimate; generates a warped image frame from the warped point cloud using a differentiable projection operation; compares the warped image frame with the context image frame to produce a self-supervised photometric loss; updates a set of estimated camera intrinsic parameters on a per-image-sequence basis using one or more gradients from the self-supervised photometric loss; and generates, based on a converged set of learned camera intrinsic parameters, a rectified image frame from an image frame captured by the camera.
    Type: Grant
    Filed: March 11, 2022
    Date of Patent: December 24, 2024
    Assignees: Toyota Research Institute, Inc., Toyota Technological Institute at Chicago
    Inventors: Vitor Guizilini, Adrien David Gaidon, Rares A. Ambrus, Igor Vasiljevic, Jiading Fang, Gregory Shakhnarovich, Matthew R. Walter
  • Publication number: 20240354991
    Abstract: Systems, methods, and other embodiments described herein relate to estimating scaled depth maps by sampling variational representations of an image using a learning model. In one embodiment, a method includes encoding data embeddings by a learning model to form conditioned latent representations using attention networks, the data embeddings including features about an image from a camera and calibration information about the camera. The method also includes computing a probability distribution of the conditioned latent representations by factoring scale priors. The method also includes sampling the probability distribution to generate variations for the data embeddings. The method also includes estimating scaled depth maps of a scene from the variations at different coordinates using the attention networks.
    Type: Application
    Filed: October 13, 2023
    Publication date: October 24, 2024
    Applicants: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki Kaisha
    Inventors: Vitor Campagnolo Guizilini, Igor Vasiljevic, Dian Chen, Adrien David Gaidon, Rares A. Ambrus