Patents by Inventor Thomas Kollar
Thomas Kollar has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Patent number: 12670211Abstract: A method for determining a complexity of a natural language query includes converting a first natural language query into executable program code, the first natural language query being a query for a first video to be answered by one or more video question answering (VideoQA) models. The method also includes generating, via a complexity model, an abstract syntax tree (AST) based on the executable program code. The method further includes determining, via the complexity model, a complexity of the first natural language query based on quantity of subtrees, from a group of subtrees, that are present in the AST.Type: GrantFiled: January 28, 2025Date of Patent: June 30, 2026Assignees: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHA, THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIVERSITYInventors: Cristobal Eyzaguirre, Igor Vasiljevic, Achal Dave, Jiajun Wu, Thomas Kollar, Juan Carlos Niebles, Pavel Tokmakov
-
Patent number: 12651365Abstract: System, methods, and other embodiments described herein relate to single-shot multi-object three-dimensional (3D) shape reconstruction and categorical six-dimensional (6D) pose and size estimation. In one embodiment, a method includes inferring a heatmap based upon a feature pyramid, where the feature pyramid is generated based upon a red green blue depth (RGB-D) image that includes objects. The method further includes sampling a 3D parameter map at locations corresponding to peaks in the heatmap, where the 3D parameter map is inferred based upon the feature pyramid, and where the locations include latent shape codes, 6D poses, and one-dimensional (1D) scales. The method further includes generating point clouds based upon the latent shape codes, the 6D poses, and the 1D scales.Type: GrantFiled: August 25, 2022Date of Patent: June 9, 2026Assignees: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki KaishaInventors: Muhammad Zubair Irshad, Thomas Kollar, Michael Laskey, Kevin Stone
-
Publication number: 20260145337Abstract: A method for training a neural network to perform 3D object manipulation is described. The method includes extracting features from each image of a synthetic stereo pair of images. The method also includes generating a low-resolution disparity image based on the features extracted from each image of the synthetic stereo pair of images. The method further includes generating, by the neural network, a feature map based on the low-resolution disparity image and one of the synthetic stereo pair of images. The method also includes manipulating an unknown object perceived from the feature map according to a perception prediction from a prediction head.Type: ApplicationFiled: January 12, 2026Publication date: May 28, 2026Applicants: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHAInventors: Thomas KOLLAR, Kevin STONE, Michael LASKEY, Mark Edward TJERSLAND
-
Patent number: 12552040Abstract: A method for training a neural network to perform 3D object manipulation is described. The method includes extracting features from each image of a synthetic stereo pair of images. The method also includes generating a low-resolution disparity image based on the features extracted from each image of the synthetic stereo pair of images. The method further includes generating, by the neural network, a feature map based on the low-resolution disparity image and one of the synthetic stereo pair of images. The method also includes manipulating an unknown object perceived from the feature map according to a perception prediction from a prediction head.Type: GrantFiled: June 13, 2022Date of Patent: February 17, 2026Assignees: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHAInventors: Thomas Kollar, Kevin Stone, Michael Laskey, Mark Edward Tjersland
-
Publication number: 20260014708Abstract: A method may include receiving an image of a robot, receiving a language instruction of a task to be performed by the robot, generating a plurality of image sequences of the robot performing the task based on the received image of the robot and the language instruction, selecting a first image sequence among the plurality of image sequences having a highest probability of performing the task, and determining a plurality of actions to be performed by a second robot to perform the task based on the first image sequence.Type: ApplicationFiled: May 13, 2025Publication date: January 15, 2026Applicants: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki Kaisha, The Trustees of Princeton University, The Regents of the University of CaliforniaInventors: Kyle Hatch, Ashwin Balakrishna, Suraj Nair, Blake Wulfe, Mikhal Itkina, Thomas Kollar, Benjamin Burchfiel, Benjamin Eysenbach, Oier Mees, Seohong Park, Sergey Levine
-
Publication number: 20250355934Abstract: A method for determining a complexity of a natural language query includes converting a first natural language query into executable program code, the first natural language query being a query for a first video to be answered by one or more video question answering (VideoQA) models. The method also includes generating, via a complexity model, an abstract syntax tree (AST) based on the executable program code. The method further includes determining, via the complexity model, a complexity of the first natural language query based on quantity of subtrees, from a group of subtrees, that are present in the AST.Type: ApplicationFiled: January 28, 2025Publication date: November 20, 2025Applicants: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHA, THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIVERSITYInventors: Cristobal EYZAGUIRRE, Igor VASILJEVIC, Achal DAVE, Jiajun WU, Thomas KOLLAR, Juan Carlos NIEBLES, Pavel TOKMAKOV
-
Patent number: 12466082Abstract: A robotic system is contemplated. The robotic system comprises a robot comprising a camera, a microphone, memory, and a controller that is configured to receive a natural language command for performing an action within a real world environment, parse the natural language command, categorize the action as being associated with guidance for performing the action, receive the guidance for performing the action, the guidance including a motion applied to at least one portion of the robot within the real world environment for performing the action, and store, in the memory, the natural language command in correlation with the motion that is applied to the at least one portion of the robot.Type: GrantFiled: June 14, 2022Date of Patent: November 11, 2025Assignees: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHAInventor: Thomas Kollar
-
Patent number: 12462801Abstract: A system capable of performing natural language understanding (NLU) on utterances including complex command structures such as sequential commands (e.g., multiple commands in a single utterance), conditional commands (e.g., commands that are only executed if a condition is satisfied), and/or repetitive commands (e.g., commands that are executed until a condition is satisfied). Audio data may be processed using automatic speech recognition (ASR) techniques to obtain text. The text may then be processed using machine learning models that are trained to parse text of incoming utterances. The models may identify complex utterance structures and may identify what command portions of an utterance go with what conditional statements. Machine learning models may also identify what data is needed to determine when the conditionals are true so the system may cause the commands to be executed (and stopped) at the appropriate times.Type: GrantFiled: August 8, 2022Date of Patent: November 4, 2025Assignee: Amazon Technologies, Inc.Inventors: Cengiz Erbas, Thomas Kollar, Avnish Sikka, Spyridon Matsoukas, Simon Peter Reavely
-
Publication number: 20250307616Abstract: A method may include receiving parameters associated with a pre-trained transformer trained on first training data, modifying an architecture of the pre-trained transformer to generate a modified transformer, the modified transformer replacing a dot-product softmax attention layer with a linear kernel squared dot product attention layer utilizing Group Normalization, receiving second training data, and training the modified transformer based on the training data.Type: ApplicationFiled: January 31, 2025Publication date: October 2, 2025Applicants: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki KaishaInventors: Jean Mercat, Igor Vasiljevic, Sedrick Keh, Achal Dave, Kushal Arora, Thomas Kollar
-
Publication number: 20250307633Abstract: A method may include receiving parameters associated with a pre-trained transformer trained on first training data, modifying an architecture of the pre-trained transformer to generate a modified transformer, the modified transformer replacing a dot-product softmax attention layer with a linear kernel dot product attention layer utilizing Group Normalization, receiving second training data, and training the modified transformer based on the training data.Type: ApplicationFiled: January 31, 2025Publication date: October 2, 2025Applicants: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki KaishaInventors: Igor Vasiljevic, Jean Mercat, Sedrick Keh, Achal Dave, Kushal Arora, Thomas Kollar
-
Publication number: 20250217996Abstract: A method for 3D object perception is described. The method includes extracting features from each image of a synthetic stereo pair of images. The method also includes generating a low-resolution disparity image based on the features extracted from each image of the synthetic stereo pair images. The method further includes predicting, by a trained neural network, a feature map based on the low-resolution disparity image and one of the synthetic stereo pair of images. The method also includes generating, by a perception prediction head, a perception prediction of a detected 3D object based on the feature map predicted by the trained neural network.Type: ApplicationFiled: March 21, 2025Publication date: July 3, 2025Applicants: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHAInventors: Thomas KOLLAR, Kevin STONE, Michael LASKEY, Mark Edward TJERSLAND
-
Patent number: 12288340Abstract: A method for 3D object perception is described. The method includes extracting features from each image of a synthetic stereo pair of images. The method also includes generating a low-resolution disparity image based on the features extracted from each image of the synthetic stereo pair images. The method further includes predicting, by a trained neural network, a feature map based on the low-resolution disparity image and one of the synthetic stereo pair of images. The method also includes generating, by a perception prediction head, a perception prediction of a detected 3D object based on the feature map predicted by the trained neural network.Type: GrantFiled: June 13, 2022Date of Patent: April 29, 2025Assignees: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHAInventors: Thomas Kollar, Kevin Stone, Michael Laskey, Mark Edward Tjersland
-
Publication number: 20240320843Abstract: Aspects of the present disclosure provide techniques for category and joint agnostic reconstruction of articulated objects. An example method includes obtaining images of an environment having objects and generating, using a trained AI encoder, first information associated with the images based at least in part on the images, the first information comprising a plurality of joint codes and a plurality of shape codes associated with the images. The method further includes generating, using a trained AI decoder, second information associated with the objects based at least in part on the plurality of joint codes and the plurality of shape codes, the second information comprising shape information, one or more joint types, and one or more joint states corresponding to at least one of the objects. The method further includes storing the second information in memory.Type: ApplicationFiled: February 14, 2024Publication date: September 26, 2024Applicants: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki Kaisha, The Board of Trustees of the Leland Stanford Junior UniversityInventors: Thomas KOLLAR, Nick HEPPERT, Muhammed Zubair IRSHAD, Rares A AMBRUS, Katherine LIU, Jeannette BOHG, Sergey ZAKHAROV
-
Publication number: 20240171724Abstract: The present disclosure provides neural fields for sparse novel view synthesis of outdoor scenes. Given just a single or a few input images from a novel scene, the disclosed technology can render new 360° views of complex unbounded outdoor scenes. This can be achieved by constructing an image-conditional triplanar representation to model the 3D surrounding from various perspectives. The disclosed technology can generalize across novel scenes and viewpoints for complex 360° outdoor scenes.Type: ApplicationFiled: October 16, 2023Publication date: May 23, 2024Applicants: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHAInventors: MUHAMMAD ZUBAIR IRSHAD, SERGEY ZAKHAROV, KATHERINE Y. LIU, VITOR GUIZILINI, THOMAS KOLLAR, ADRIEN D. GAIDON, RARES A. AMBRUS
-
Publication number: 20230401721Abstract: A method for 3D object perception is described. The method includes extracting features from each image of a synthetic stereo pair of images. The method also includes generating a low-resolution disparity image based on the features extracted from each image of the synthetic stereo pair images. The method further includes predicting, by a trained neural network, a feature map based on the low-resolution disparity image and one of the synthetic stereo pair of images. The method also includes generating, by a perception prediction head, a perception prediction of a detected 3D object based on the feature map predicted by the trained neural network.Type: ApplicationFiled: June 13, 2022Publication date: December 14, 2023Applicants: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHAInventors: Thomas KOLLAR, Kevin STONE, Michael LASKEY, Mark Edward TJERSLAND
-
Publication number: 20230398696Abstract: A robotic system is contemplated. The robotic system comprises a robot comprising a camera, a microphone, memory, and a controller that is configured to receive a natural language command for performing an action within a real world environment, parse the natural language command, categorize the action as being associated with guidance for performing the action, receive the guidance for performing the action, the guidance including a motion applied to at least one portion of the robot within the real world environment for performing the action, and store, in the memory, the natural language command in correlation with the motion that is applied to the at least one portion of the robot.Type: ApplicationFiled: June 14, 2022Publication date: December 14, 2023Applicants: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki KaishaInventor: Thomas Kollar
-
Publication number: 20230398692Abstract: A method for training a neural network to perform 3D object manipulation is described. The method includes extracting features from each image of a synthetic stereo pair of images. The method also includes generating a low-resolution disparity image based on the features extracted from each image of the synthetic stereo pair of images. The method further includes generating, by the neural network, a feature map based on the low-resolution disparity image and one of the synthetic stereo pair of images. The method also includes manipulating an unknown object perceived from the feature map according to a perception prediction from a prediction head.Type: ApplicationFiled: June 13, 2022Publication date: December 14, 2023Applicants: TOYOTA RESEARCH INSTITUTE, INC., TOYOTA JIDOSHA KABUSHIKI KAISHAInventors: Thomas KOLLAR, Kevin STONE, Michael LASKEY, Mark Edward TJERSLAND
-
Publication number: 20230077856Abstract: System, methods, and other embodiments described herein relate to single-shot multi-object three-dimensional (3D) shape reconstruction and categorical six-dimensional (6D) pose and size estimation. In one embodiment, a method includes inferring a heatmap based upon a feature pyramid, where the feature pyramid is generated based upon a red green blue depth (RGB-D) image that includes objects. The method further includes sampling a 3D parameter map at locations corresponding to peaks in the heatmap, where the 3D parameter map is inferred based upon the feature pyramid, and where the locations include latent shape codes, 6D poses, and one-dimensional (1D) scales. The method further includes generating point clouds based upon the latent shape codes, the 6D poses, and the 1D scales.Type: ApplicationFiled: August 25, 2022Publication date: March 16, 2023Applicants: Toyota Research Institute, Inc., Toyota Jidosha Kabushiki KaishaInventors: Muhammad Zubair Irshad, Thomas Kollar, Michael Laskey, Kevin Stone
-
Publication number: 20230032575Abstract: A system capable of performing natural language understanding (NLU) on utterances including complex command structures such as sequential commands (e.g., multiple commands in a single utterance), conditional commands (e.g., commands that are only executed if a condition is satisfied), and/or repetitive commands (e.g., commands that are executed until a condition is satisfied). Audio data may be processed using automatic speech recognition (ASR) techniques to obtain text. The text may then be processed using machine learning models that are trained to parse text of incoming utterances. The models may identify complex utterance structures and may identify what command portions of an utterance go with what conditional statements. Machine learning models may also identify what data is needed to determine when the conditionals are true so the system may cause the commands to be executed (and stopped) at the appropriate times.Type: ApplicationFiled: August 8, 2022Publication date: February 2, 2023Inventors: Cengiz Erbas, Thomas Kollar, Avnish Sikka, Spyridon Matsoukas, Simon Peter Reavely
-
Patent number: 11410646Abstract: A system capable of performing natural language understanding (NLU) on utterances including complex command structures such as sequential commands (e.g., multiple commands in a single utterance), conditional commands (e.g., commands that are only executed if a condition is satisfied), and/or repetitive commands (e.g., commands that are executed until a condition is satisfied). Audio data may be processed using automatic speech recognition (ASR) techniques to obtain text. The text may then be processed using machine learning models that are trained to parse text of incoming utterances. The models may identify complex utterance structures and may identify what command portions of an utterance go with what conditional statements. Machine learning models may also identify what data is needed to determine when the conditionals are true so the system may cause the commands to be executed (and stopped) at the appropriate times.Type: GrantFiled: March 28, 2019Date of Patent: August 9, 2022Assignee: Amazon Technologies, Inc.Inventors: Cengiz Erbas, Thomas Kollar, Avnish Sikka, Spyridon Matsoukas, Simon Peter Reavely