Patents by Inventor Sameer Khurana

Sameer Khurana has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20250390682
    Abstract: Systems, methods, software, and devices are disclosed herein process context data to encode one or more semantic elements of a desired audio composition in a semantic token sequence, process the semantic token sequence to encode one or more structural elements of the desired audio composition in a structural token sequence disentangled from the semantic token sequence, and process the structural token sequence to encode one or more audio signal elements of the desired audio composition in an audio signal token sequence disentangled from the structural token sequence. The semantic token sequence, the structural token sequence, and the audio signal token sequence may then be processed to generate at least a portion of the desired audio composition.
    Type: Application
    Filed: November 13, 2024
    Publication date: December 25, 2025
    Applicant: Mitsubishi Electric Research Laboratories, Inc.
    Inventors: Sameer Khurana, Jonathan Le Roux, Gordon Wichern, Chiori Hori, François G Germain, Janek Ebbers, Kohei Saijo, Amir Hussein
  • Publication number: 20250355419
    Abstract: A robotic controller for controlling a robot according to a sequence of robotic actions. comprises an input interface configured to receive a plurality of multimodal inputs each specifying instructions for performing a task in a different modality including audio, video, and a text modality. The controller also comprises a multimodal large language model, an action sequence decoder, and a controller. The multimodal LLM includes a multimodal LLM encoder and an LLM decoder. The multimodal LLM encoder is trained with machine learning to transform the multimodal instructions into encodings and the LLM decoder is configured to decode the encodings into a sequence of robotic instructions. The action sequence decoder is trained with machine learning to transform the sequence of robotic instructions into a sequence of actions using a library of robotic skills. The controller is configured to control a robot according to the sequence of actions.
    Type: Application
    Filed: July 16, 2024
    Publication date: November 20, 2025
    Applicant: Mitsubishi Electric Research Laboratories, Inc.
    Inventors: Chiori Hori, Motonari Kambara, Devesh Jha, Diego Romeres, Siddarth Jain, Radu Ioan Corcodel, Kei Ota, Jonathan Le Roux, Sameer Khurana
  • Publication number: 20250353175
    Abstract: A robotic controller for controlling a robot according to a sequence of robotic actions. comprises an input interface to receive multimodal inputs specifying instructions for performing a task in audio, video, and a text modality. The controller transforms the multimodal instructions into encodings using a large language model (LLM) encoder and decodes the encodings into a first sequence of robotic instructions and a robot action description of the actions using an LLM decoder. Human feedback input is received corresponding to at least one action in the first sequence of actions and the controller encodes the feedback input with the robot action description. The controller feeds the encoded data along with multimodal features generated from the encodings into the LLM decoder to generate a corrected sequence of actions. The controller is configured to control a robot according to the corrected sequence of actions.
    Type: Application
    Filed: March 4, 2025
    Publication date: November 20, 2025
    Inventors: Chiori Hori, Motonari Kambara, Sameer Khurana, Kei Ota, Siddarth Jain, Radu Ioan Corcodel, Devesh Jha, Diego Romeres, Jonathan Le Roux
  • Publication number: 20250292760
    Abstract: An audio system for synthesizing audio sounds having a desired audio trait executes an autoregressive generative audio transformer trained for generating the audio by processing inputs with multiple layers employing multi-head attention, and uses directional inference-time intervention (ITI) to push at least some outputs of at least some heads of the multi-head attention into a direction predetermined for the desired audio trait.
    Type: Application
    Filed: March 15, 2024
    Publication date: September 18, 2025
    Applicant: Mitsubishi Electric Research Laboratories, Inc.
    Inventors: Gordon Wichern, Junghyun Koo, François G Germain, Sameer Khurana, Jonathan Le Roux