Patents by Inventor Michael S. Ryoo

Michael S. Ryoo has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20260093993
    Abstract: Embodiments described herein are directed to training a large action model (LAM) simulator framework. The LAM framework receives a content dataset associated with a task. This dataset includes the task name and at least one user command parameter. The LAM simulator framework identifies an abstract task based on the task name and determines the available tools for the task. It then generates a user command for an artificial intelligence agent, instructing it to complete the task using the abstract task and user command parameters. The AI agent is trained to execute the task over multiple iterations. An iteration involves creating a conversation data object from the user command, available tools, and prior conversation history, generating an action plan using a generative language model and the conversation data object, executing the actions in the plan using the environment and tools, and evaluating the actions.
    Type: Application
    Filed: January 28, 2025
    Publication date: April 2, 2026
    Inventors: Thai Hoang, Shirley Kokane, Jianguo Zhang, Tian Lan, Zuxin Liu, Ming Zhu, Jake Grigsby, Michael S. Ryoo, Shelby Heinecke, Caiming Xiong, Huan Wang, Juan Carlos Niebles Duque, Silvio Savarese
  • Publication number: 20260080681
    Abstract: Embodiments described herein provide a vision-language neural network framework that outputs a text response to a user text query relating to the media content of the video input. Specifically, the vision-language neural network may comprise (1) a vision encoder (ViT) transforming each frame input from the video input into a set of tokens, (2) a frame-level tokenizer to reduce the number of tokens, (3) a temporal encoder to build video-level token representations, and (4) an autoregressive LLM generating a text output based on such video tokens and text prompt tokens.
    Type: Application
    Filed: January 30, 2025
    Publication date: March 19, 2026
    Inventors: Michael S. Ryoo, Honglu Zhou, Shrikant Kendre, Can Qin, Le Xue, Manli Shu, Silvio Savarese, Ran Xu, Caiming Xiong, Juan Carlos Niebles Duque
  • Publication number: 20260044993
    Abstract: Embodiments described herein provide a generation model comprising a video-specific variational auto-encoder (VAE) for effective compression of video pixel information with reduced spatial and temporal dimensions and a video diffusion transformer (vDiT) to generate latent representations of frames. Specifically, the VAE may, instead of encoding each frame independently, incorporate both temporal and spatial compression. This significantly decreases the token length, improves the computational cost of training and inference, and facilitates the generation of long videos. The encoded training video, in the form of latent representations from a VAE encoder may then be passed to the vDiT to reconstruct the latent representations during training. The trained vDiT may then generate latent representations of a video in response to a text input, and the latent representations may be converted to a video output by a VAE decoder.
    Type: Application
    Filed: January 2, 2025
    Publication date: February 12, 2026
    Inventors: Can Qin, Krithika Ramakrishnan, Congying Xia, Yihao Feng, Michael S. Ryoo, Lifu Tu, Zeyuan Chen, Ran Xu, Caiming Xiong