Patents by Inventor Daniel Matthew Cer

Daniel Matthew Cer has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20260203612
    Abstract: Systems and methods for prompt tuning can utilize previously-learned prompts for the initialization of tuning for prompts on different tasks that may differ from the task associated with the previously-learned prompt. The prompt being utilized for initialization can be a generic prompt and/or may be a prompt selected based on a determined similarity between two or more task embeddings.
    Type: Application
    Filed: March 11, 2026
    Publication date: July 16, 2026
    Inventors: Tu Thanh Vu, Daniel Matthew Cer, Noah Constant, Brian David Lester, Rami Al-Rfou
  • Patent number: 12619886
    Abstract: Systems and methods for prompt tuning can utilize previously-learned prompts for the initialization of tuning for prompts on different tasks that may differ from the task associated with the previously-learned prompt. The prompt being utilized for initialization can be a generic prompt and/or may be a prompt selected based on a determined similarity between two or more task embeddings.
    Type: Grant
    Filed: July 13, 2022
    Date of Patent: May 5, 2026
    Assignee: GOOGLE LLC
    Inventors: Tu Thanh Vu, Daniel Matthew Cer, Noah Constant, Brian David Lester, Rami Al-Rfou
  • Publication number: 20260044714
    Abstract: Aspects of the disclosure relate to pre-format embeddings used for model training, and in particular, training of dual encoder LLMs. For instance, a dataset batch of pairs of input text and one or more classification labels may be accessed. A unique identifier to each pair of input text and one or more classification labels may be assigned. The pairs of input text and one or more classification labels and assigned unique identifiers may be used to train a model to assign classification labels to textual inputs.
    Type: Application
    Filed: August 12, 2024
    Publication date: February 12, 2026
    Inventors: Daniel Matthew Cer, Jinhyuk Lee, Xiaoqi Ren, Wen Ding, Blair Yuxin Chen
  • Publication number: 20250265417
    Abstract: The technology employs soft knowledge prompts (KPs) to inject relevant world knowledge into language models. This includes training KPs via self-supervised learning on data from one or more knowledge bases. KPs are task independent and can function as an external memory of the language models. KPs may be entity-centric, meaning that each prompt primarily encodes information about one entity from a given knowledge base. A method includes identifying a KP in response to a received input text, concatenating that KP to a sequence of word embeddings of the input text, applying the concatenated information to a trained language model, predicting an object entity name, computing a cross-entropy loss, and updating the identified KP based on the computed cross-entropy loss.
    Type: Application
    Filed: May 5, 2025
    Publication date: August 21, 2025
    Inventors: Siamak Shakeri, Cicero Nogueira dos Santos, Daniel Matthew Cer, Zhe Dong, Jianmo Ni, Yun-Hsuan Sung, John Nham
  • Patent number: 12321706
    Abstract: The technology employs soft knowledge prompts (KPs) to inject relevant world knowledge into language models. This includes training KPs via self-supervised learning on data from one or more knowledge bases. KPs are task independent and can function as an external memory of the language models. KPs may be entity-centric, meaning that each prompt primarily encodes information about one entity from a given knowledge base. A method includes identifying a KP in response to a received input text, concatenating that KP to a sequence of word embeddings of the input text, applying the concatenated information to a trained language model, predicting an object entity name, computing a cross-entropy loss, and updating the identified KP based on the computed cross-entropy loss.
    Type: Grant
    Filed: February 9, 2023
    Date of Patent: June 3, 2025
    Assignee: Google LLC
    Inventors: Siamak Shakeri, Cicero Nogueira dos Santos, Daniel Matthew Cer, Zhe Dong, Jianmo Ni, Yun-Hsuan Sung, John Nham
  • Publication number: 20250045316
    Abstract: An example method includes providing, to a sequence model (i) a plurality of few-shot prompts, wherein each prompt comprises a demonstration passage, a demonstration task, and a demonstration query, wherein the demonstration task describes a type of retrieval, and wherein the demonstration query is relevant to the demonstration task, and (ii) a plurality of passages sampled from a corpus of passages. The method also includes receiving, from the sequence model and for the plurality of passages and based on the plurality of few-shot prompts, a respective plurality of predicted task-query pairs, the sequence model having been prompted to predict a task based on an input passage, and predict an output query relevant to the predicted task. The method further includes generating a synthetic training dataset comprising the plurality of passages and the respective plurality of predicted task-query pairs. The method also includes providing the synthetic training dataset.
    Type: Application
    Filed: July 30, 2024
    Publication date: February 6, 2025
    Inventors: Jinhyuk Lee, Zhuyun Dai, Xiaoqi Ren, Iftekhar Naim, Yi Luan, Blair Yuxin Chen, Siddhartha Reddy Jonnalagadda, Ming-Wei Chang, Daniel Matthew Cer, Gustavo Adolfo Hernandez Abrego, Jeremy Robert Cole, Colin Hearne Evans, Yuzhe Zhao, Pranay Bhatia, Rajvi Kapadia, Riham Hassan Abdel-Moneim Mansour, Raphael Dominik Hoffman, Simon Kunio Tokumine, Scott Bradley Huffman, Stephen Zachary Karukas, Michael Yiupun Kwong, Shu Zheng, Yan Qiao, Lukas Rutishauser, Anand Rajan Iyer
  • Publication number: 20240273294
    Abstract: The technology employs soft knowledge prompts (KPs) to inject relevant world knowledge into language models. This includes training KPs via self-supervised learning on data from one or more knowledge bases. KPs are task independent and can function as an external memory of the language models. KPs may be entity-centric, meaning that each prompt primarily encodes information about one entity from a given knowledge base. A method includes identifying a KP in response to a received input text, concatenating that KP to a sequence of word embeddings of the input text, applying the concatenated information to a trained language model, predicting an object entity name, computing a cross-entropy loss, and updating the identified KP based on the computed cross-entropy loss.
    Type: Application
    Filed: February 9, 2023
    Publication date: August 15, 2024
    Inventors: Siamak Shakeri, Cicero Nogueira dos Santos, Daniel Matthew Cer, Zhe Dong, Jianmo Ni, Yun-Hsuan Sung, John Nham
  • Publication number: 20240020546
    Abstract: Systems and methods for prompt tuning can utilize previously-learned prompts for the initialization of tuning for prompts on different tasks that may differ from the task associated with the previously-learned prompt. The prompt being utilized for initialization can be a generic prompt and/or may be a prompt selected based on a determined similarity between two or more task embeddings.
    Type: Application
    Filed: July 13, 2022
    Publication date: January 18, 2024
    Inventors: Tu Thanh Vu, Daniel Matthew Cer, Noah Constant, Brian David Lester, Rami Al-Rfou
  • Patent number: 11769011
    Abstract: The present disclosure provides a novel sentence-level representation learning method Conditional Masked Language Modeling (CMLM) for training on large scale unlabeled corpora. CMLM outperforms the previous state-of-the-art English sentence embedding models, including those trained with (semi-)supervised signals. For multilingual representations learning, it is shown that co-training CMLM with bitext retrieval and cross-lingual natural language inference (NL) fine-tuning achieves state-of-the-art performance. It is also shown that multilingual representations have the same language bias and principal component removal (PCR) can eliminate the bias by separating language identity information from semantics.
    Type: Grant
    Filed: December 18, 2020
    Date of Patent: September 26, 2023
    Assignee: GOOGLE LLC
    Inventors: Yinfei Yang, Ziyi Yang, Daniel Matthew Cer
  • Publication number: 20220198144
    Abstract: The present disclosure provides a novel sentence-level representation learning method Conditional Masked Language Modeling (CMLM) for training on large scale unlabeled corpora. CMLM outperforms the previous state-of-the-art English sentence embedding models, including those trained with (semi-)supervised signals. For multilingual representations learning, it is shown that co-training CMLM with bitext retrieval and cross-lingual NLI fine-tuning achieves state-of-the-art performance. It is also shown that multilingual representations have the same language bias and principal component removal (PCR) can eliminate the bias by separating language identity information from semantics.
    Type: Application
    Filed: December 18, 2020
    Publication date: June 23, 2022
    Inventors: Yinfei Yang, Ziyi Yang, Daniel Matthew Cer