Patents by Inventor Thomas Mesnard

Thomas Mesnard has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20250299055
    Abstract: Aspects of the disclosure are directed to using reinforcement learning to train one or more machine learning models based on reward data that is model generated. The reward data is generated by a generative model, such as a large language model, in response to a prompt to provide respective reward scores for model-generated responses to a task. Since generating preference labels and training of a reward model can be bypassed here, the machine learning models can be trained using reinforcement learning with less processing cost and memory usage.
    Type: Application
    Filed: March 19, 2024
    Publication date: September 25, 2025
    Inventors: Samrat Phatale, Harrison Lee, Hassan Mansoor, Kellie Lu, Thomas Mesnard, Johan Ferret, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Kumar Rastogi, Sushant Prakash, Mo Azar, Zhaohan Daniel Guo, Andrea Michi, Nicolas Perez Nieves, Marco Selvi
  • Publication number: 20250068919
    Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training a neural network used to select actions to be performed by an agent interacting with an environment. Implementations of the method model unpredictable aspects of the future, using hindsight. They use this information to disentangle inherently unpredictable, aleatoric variation, from epistemic uncertainty that arises from lack of knowledge of the environment. They then use the epistemic uncertainty, which relates to in principle predictable aspects of the environment, as a source of intrinsic reward to drive curiosity, i.e. exploration of the environment by the agent.
    Type: Application
    Filed: August 25, 2023
    Publication date: February 27, 2025
    Inventors: Daniel Jarrett, Corentin Tallec, Florent Altché, Thomas Mesnard, Remi Munos, Michal Valko
  • Publication number: 20240256883
    Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training a neural network used to select actions to be performed by an agent interacting with an environment. Implementations of the system can take into account a level of luck in the environment, and hence whilst learning can account for outcomes that were caused by external factors as well as those dependent on the actions of the agent.
    Type: Application
    Filed: January 26, 2024
    Publication date: August 1, 2024
    Inventors: Thomas Mesnard, Remi Munos, Alaa Saade, Yunhao Tang, Mark Daniel Rowland, Theophane Guillaume Weber, Wenqi Chen