Patents by Inventor Thomas Kollar

Thomas Kollar has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Patent number: 12670211
    Abstract: A method for determining a complexity of a natural language query includes converting a first natural language query into executable program code, the first natural language query being a query for a first video to be answered by one or more video question answering (VideoQA) models. The method also includes generating, via a complexity model, an abstract syntax tree (AST) based on the executable program code. The method further includes determining, via the complexity model, a complexity of the first natural language query based on quantity of subtrees, from a group of subtrees, that are present in the AST.
    Type: Grant
    Filed: January 28, 2025
    Date of Patent: June 30, 2026
    Assignees: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHA, THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIVERSITY
    Inventors: Cristobal Eyzaguirre, Igor Vasiljevic, Achal Dave, Jiajun Wu, Thomas Kollar, Juan Carlos Niebles, Pavel Tokmakov
  • Patent number: 12651365
    Abstract: System, methods, and other embodiments described herein relate to single-shot multi-object three-dimensional (3D) shape reconstruction and categorical six-dimensional (6D) pose and size estimation. In one embodiment, a method includes inferring a heatmap based upon a feature pyramid, where the feature pyramid is generated based upon a red green blue depth (RGB-D) image that includes objects. The method further includes sampling a 3D parameter map at locations corresponding to peaks in the heatmap, where the 3D parameter map is inferred based upon the feature pyramid, and where the locations include latent shape codes, 6D poses, and one-dimensional (1D) scales. The method further includes generating point clouds based upon the latent shape codes, the 6D poses, and the 1D scales.
    Type: Grant
    Filed: August 25, 2022
    Date of Patent: June 9, 2026
    Assignees: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki Kaisha
    Inventors: Muhammad Zubair Irshad, Thomas Kollar, Michael Laskey, Kevin Stone
  • Publication number: 20260145337
    Abstract: A method for training a neural network to perform 3D object manipulation is described. The method includes extracting features from each image of a synthetic stereo pair of images. The method also includes generating a low-resolution disparity image based on the features extracted from each image of the synthetic stereo pair of images. The method further includes generating, by the neural network, a feature map based on the low-resolution disparity image and one of the synthetic stereo pair of images. The method also includes manipulating an unknown object perceived from the feature map according to a perception prediction from a prediction head.
    Type: Application
    Filed: January 12, 2026
    Publication date: May 28, 2026
    Applicants: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHA
    Inventors: Thomas KOLLAR, Kevin STONE, Michael LASKEY, Mark Edward TJERSLAND
  • Patent number: 12552040
    Abstract: A method for training a neural network to perform 3D object manipulation is described. The method includes extracting features from each image of a synthetic stereo pair of images. The method also includes generating a low-resolution disparity image based on the features extracted from each image of the synthetic stereo pair of images. The method further includes generating, by the neural network, a feature map based on the low-resolution disparity image and one of the synthetic stereo pair of images. The method also includes manipulating an unknown object perceived from the feature map according to a perception prediction from a prediction head.
    Type: Grant
    Filed: June 13, 2022
    Date of Patent: February 17, 2026
    Assignees: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHA
    Inventors: Thomas Kollar, Kevin Stone, Michael Laskey, Mark Edward Tjersland
  • Publication number: 20260014708
    Abstract: A method may include receiving an image of a robot, receiving a language instruction of a task to be performed by the robot, generating a plurality of image sequences of the robot performing the task based on the received image of the robot and the language instruction, selecting a first image sequence among the plurality of image sequences having a highest probability of performing the task, and determining a plurality of actions to be performed by a second robot to perform the task based on the first image sequence.
    Type: Application
    Filed: May 13, 2025
    Publication date: January 15, 2026
    Applicants: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki Kaisha, The Trustees of Princeton University, The Regents of the University of California
    Inventors: Kyle Hatch, Ashwin Balakrishna, Suraj Nair, Blake Wulfe, Mikhal Itkina, Thomas Kollar, Benjamin Burchfiel, Benjamin Eysenbach, Oier Mees, Seohong Park, Sergey Levine
  • Publication number: 20250355934
    Abstract: A method for determining a complexity of a natural language query includes converting a first natural language query into executable program code, the first natural language query being a query for a first video to be answered by one or more video question answering (VideoQA) models. The method also includes generating, via a complexity model, an abstract syntax tree (AST) based on the executable program code. The method further includes determining, via the complexity model, a complexity of the first natural language query based on quantity of subtrees, from a group of subtrees, that are present in the AST.
    Type: Application
    Filed: January 28, 2025
    Publication date: November 20, 2025
    Applicants: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHA, THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIVERSITY
    Inventors: Cristobal EYZAGUIRRE, Igor VASILJEVIC, Achal DAVE, Jiajun WU, Thomas KOLLAR, Juan Carlos NIEBLES, Pavel TOKMAKOV
  • Patent number: 12466082
    Abstract: A robotic system is contemplated. The robotic system comprises a robot comprising a camera, a microphone, memory, and a controller that is configured to receive a natural language command for performing an action within a real world environment, parse the natural language command, categorize the action as being associated with guidance for performing the action, receive the guidance for performing the action, the guidance including a motion applied to at least one portion of the robot within the real world environment for performing the action, and store, in the memory, the natural language command in correlation with the motion that is applied to the at least one portion of the robot.
    Type: Grant
    Filed: June 14, 2022
    Date of Patent: November 11, 2025
    Assignees: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHA
    Inventor: Thomas Kollar
  • Patent number: 12462801
    Abstract: A system capable of performing natural language understanding (NLU) on utterances including complex command structures such as sequential commands (e.g., multiple commands in a single utterance), conditional commands (e.g., commands that are only executed if a condition is satisfied), and/or repetitive commands (e.g., commands that are executed until a condition is satisfied). Audio data may be processed using automatic speech recognition (ASR) techniques to obtain text. The text may then be processed using machine learning models that are trained to parse text of incoming utterances. The models may identify complex utterance structures and may identify what command portions of an utterance go with what conditional statements. Machine learning models may also identify what data is needed to determine when the conditionals are true so the system may cause the commands to be executed (and stopped) at the appropriate times.
    Type: Grant
    Filed: August 8, 2022
    Date of Patent: November 4, 2025
    Assignee: Amazon Technologies, Inc.
    Inventors: Cengiz Erbas, Thomas Kollar, Avnish Sikka, Spyridon Matsoukas, Simon Peter Reavely
  • Publication number: 20250307616
    Abstract: A method may include receiving parameters associated with a pre-trained transformer trained on first training data, modifying an architecture of the pre-trained transformer to generate a modified transformer, the modified transformer replacing a dot-product softmax attention layer with a linear kernel squared dot product attention layer utilizing Group Normalization, receiving second training data, and training the modified transformer based on the training data.
    Type: Application
    Filed: January 31, 2025
    Publication date: October 2, 2025
    Applicants: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki Kaisha
    Inventors: Jean Mercat, Igor Vasiljevic, Sedrick Keh, Achal Dave, Kushal Arora, Thomas Kollar
  • Publication number: 20250307633
    Abstract: A method may include receiving parameters associated with a pre-trained transformer trained on first training data, modifying an architecture of the pre-trained transformer to generate a modified transformer, the modified transformer replacing a dot-product softmax attention layer with a linear kernel dot product attention layer utilizing Group Normalization, receiving second training data, and training the modified transformer based on the training data.
    Type: Application
    Filed: January 31, 2025
    Publication date: October 2, 2025
    Applicants: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki Kaisha
    Inventors: Igor Vasiljevic, Jean Mercat, Sedrick Keh, Achal Dave, Kushal Arora, Thomas Kollar
  • Publication number: 20250217996
    Abstract: A method for 3D object perception is described. The method includes extracting features from each image of a synthetic stereo pair of images. The method also includes generating a low-resolution disparity image based on the features extracted from each image of the synthetic stereo pair images. The method further includes predicting, by a trained neural network, a feature map based on the low-resolution disparity image and one of the synthetic stereo pair of images. The method also includes generating, by a perception prediction head, a perception prediction of a detected 3D object based on the feature map predicted by the trained neural network.
    Type: Application
    Filed: March 21, 2025
    Publication date: July 3, 2025
    Applicants: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHA
    Inventors: Thomas KOLLAR, Kevin STONE, Michael LASKEY, Mark Edward TJERSLAND
  • Patent number: 12288340
    Abstract: A method for 3D object perception is described. The method includes extracting features from each image of a synthetic stereo pair of images. The method also includes generating a low-resolution disparity image based on the features extracted from each image of the synthetic stereo pair images. The method further includes predicting, by a trained neural network, a feature map based on the low-resolution disparity image and one of the synthetic stereo pair of images. The method also includes generating, by a perception prediction head, a perception prediction of a detected 3D object based on the feature map predicted by the trained neural network.
    Type: Grant
    Filed: June 13, 2022
    Date of Patent: April 29, 2025
    Assignees: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHA
    Inventors: Thomas Kollar, Kevin Stone, Michael Laskey, Mark Edward Tjersland
  • Publication number: 20240320843
    Abstract: Aspects of the present disclosure provide techniques for category and joint agnostic reconstruction of articulated objects. An example method includes obtaining images of an environment having objects and generating, using a trained AI encoder, first information associated with the images based at least in part on the images, the first information comprising a plurality of joint codes and a plurality of shape codes associated with the images. The method further includes generating, using a trained AI decoder, second information associated with the objects based at least in part on the plurality of joint codes and the plurality of shape codes, the second information comprising shape information, one or more joint types, and one or more joint states corresponding to at least one of the objects. The method further includes storing the second information in memory.
    Type: Application
    Filed: February 14, 2024
    Publication date: September 26, 2024
    Applicants: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki Kaisha, The Board of Trustees of the Leland Stanford Junior University
    Inventors: Thomas KOLLAR, Nick HEPPERT, Muhammed Zubair IRSHAD, Rares A AMBRUS, Katherine LIU, Jeannette BOHG, Sergey ZAKHAROV
  • Publication number: 20240171724
    Abstract: The present disclosure provides neural fields for sparse novel view synthesis of outdoor scenes. Given just a single or a few input images from a novel scene, the disclosed technology can render new 360° views of complex unbounded outdoor scenes. This can be achieved by constructing an image-conditional triplanar representation to model the 3D surrounding from various perspectives. The disclosed technology can generalize across novel scenes and viewpoints for complex 360° outdoor scenes.
    Type: Application
    Filed: October 16, 2023
    Publication date: May 23, 2024
    Applicants: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHA
    Inventors: MUHAMMAD ZUBAIR IRSHAD, SERGEY ZAKHAROV, KATHERINE Y. LIU, VITOR GUIZILINI, THOMAS KOLLAR, ADRIEN D. GAIDON, RARES A. AMBRUS
  • Publication number: 20230401721
    Abstract: A method for 3D object perception is described. The method includes extracting features from each image of a synthetic stereo pair of images. The method also includes generating a low-resolution disparity image based on the features extracted from each image of the synthetic stereo pair images. The method further includes predicting, by a trained neural network, a feature map based on the low-resolution disparity image and one of the synthetic stereo pair of images. The method also includes generating, by a perception prediction head, a perception prediction of a detected 3D object based on the feature map predicted by the trained neural network.
    Type: Application
    Filed: June 13, 2022
    Publication date: December 14, 2023
    Applicants: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHA
    Inventors: Thomas KOLLAR, Kevin STONE, Michael LASKEY, Mark Edward TJERSLAND
  • Publication number: 20230398696
    Abstract: A robotic system is contemplated. The robotic system comprises a robot comprising a camera, a microphone, memory, and a controller that is configured to receive a natural language command for performing an action within a real world environment, parse the natural language command, categorize the action as being associated with guidance for performing the action, receive the guidance for performing the action, the guidance including a motion applied to at least one portion of the robot within the real world environment for performing the action, and store, in the memory, the natural language command in correlation with the motion that is applied to the at least one portion of the robot.
    Type: Application
    Filed: June 14, 2022
    Publication date: December 14, 2023
    Applicants: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki Kaisha
    Inventor: Thomas Kollar
  • Publication number: 20230398692
    Abstract: A method for training a neural network to perform 3D object manipulation is described. The method includes extracting features from each image of a synthetic stereo pair of images. The method also includes generating a low-resolution disparity image based on the features extracted from each image of the synthetic stereo pair of images. The method further includes generating, by the neural network, a feature map based on the low-resolution disparity image and one of the synthetic stereo pair of images. The method also includes manipulating an unknown object perceived from the feature map according to a perception prediction from a prediction head.
    Type: Application
    Filed: June 13, 2022
    Publication date: December 14, 2023
    Applicants: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHA
    Inventors: Thomas KOLLAR, Kevin STONE, Michael LASKEY, Mark Edward TJERSLAND
  • Publication number: 20230077856
    Abstract: System, methods, and other embodiments described herein relate to single-shot multi-object three-dimensional (3D) shape reconstruction and categorical six-dimensional (6D) pose and size estimation. In one embodiment, a method includes inferring a heatmap based upon a feature pyramid, where the feature pyramid is generated based upon a red green blue depth (RGB-D) image that includes objects. The method further includes sampling a 3D parameter map at locations corresponding to peaks in the heatmap, where the 3D parameter map is inferred based upon the feature pyramid, and where the locations include latent shape codes, 6D poses, and one-dimensional (1D) scales. The method further includes generating point clouds based upon the latent shape codes, the 6D poses, and the 1D scales.
    Type: Application
    Filed: August 25, 2022
    Publication date: March 16, 2023
    Applicants: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki Kaisha
    Inventors: Muhammad Zubair Irshad, Thomas Kollar, Michael Laskey, Kevin Stone
  • Publication number: 20230032575
    Abstract: A system capable of performing natural language understanding (NLU) on utterances including complex command structures such as sequential commands (e.g., multiple commands in a single utterance), conditional commands (e.g., commands that are only executed if a condition is satisfied), and/or repetitive commands (e.g., commands that are executed until a condition is satisfied). Audio data may be processed using automatic speech recognition (ASR) techniques to obtain text. The text may then be processed using machine learning models that are trained to parse text of incoming utterances. The models may identify complex utterance structures and may identify what command portions of an utterance go with what conditional statements. Machine learning models may also identify what data is needed to determine when the conditionals are true so the system may cause the commands to be executed (and stopped) at the appropriate times.
    Type: Application
    Filed: August 8, 2022
    Publication date: February 2, 2023
    Inventors: Cengiz Erbas, Thomas Kollar, Avnish Sikka, Spyridon Matsoukas, Simon Peter Reavely
  • Patent number: 11410646
    Abstract: A system capable of performing natural language understanding (NLU) on utterances including complex command structures such as sequential commands (e.g., multiple commands in a single utterance), conditional commands (e.g., commands that are only executed if a condition is satisfied), and/or repetitive commands (e.g., commands that are executed until a condition is satisfied). Audio data may be processed using automatic speech recognition (ASR) techniques to obtain text. The text may then be processed using machine learning models that are trained to parse text of incoming utterances. The models may identify complex utterance structures and may identify what command portions of an utterance go with what conditional statements. Machine learning models may also identify what data is needed to determine when the conditionals are true so the system may cause the commands to be executed (and stopped) at the appropriate times.
    Type: Grant
    Filed: March 28, 2019
    Date of Patent: August 9, 2022
    Assignee: Amazon Technologies, Inc.
    Inventors: Cengiz Erbas, Thomas Kollar, Avnish Sikka, Spyridon Matsoukas, Simon Peter Reavely