Patents by Inventor Khalid Salama

Khalid Salama has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20260203363
    Abstract: Provided are computer-implemented systems and methods for responding to visual queries using both a visual search engine and a vision language model (VLM). In particular, aspects of the present disclosure can improve the performance of a VLM at generating a response to a visual queries by supplementing the visual query with one or more annotations generated by or using the visual search engine.
    Type: Application
    Filed: January 10, 2025
    Publication date: July 16, 2026
    Inventors: Fabio Luca Sulser, Susan Qi Xu, Vikas Bahirwani, Bhanu Prakash Reddy Guda, Lin Li, Khalid Salama, Manuel Tragut, Ágoston Weisz, Andrea Colaco
  • Publication number: 20260162325
    Abstract: Implementations relate to systems and methods for controlling generative model processing of image data containing specific objects or classifications. User requests including an image are processed, identifying whether the image includes object(s) with restricted features or classifications. When such object(s) are identified, the image and/or textual description(s) of the image are modified to omit restricted objects. An input prompt for a generative model can be generated based on the modified image and/or the modified textual description(s) of the image, that omit the restricted object(s), ensuring the generated content that is responsive to the request is not generated based on the restricted object(s) and/or omits any mention of the restricted object(s). Implementations maintain data security and/or provide computational efficiencies by selectively excluding certain information from generative model processing.
    Type: Application
    Filed: December 5, 2024
    Publication date: June 11, 2026
    Inventors: Agoston Weisz, Khalid Salama, Diana Avram
  • Patent number: 12651122
    Abstract: Implementations relate to handling visual content across a multi-turn dialog. A user input that includes natural language content and visual content is received during the dialog. If the visual content is being received for the first time in the dialog, the visual content is processed to generate a corresponding tokenized representation of the visual content. The corresponding tokenized representation can be cached in a database in association with the dialog, or in association with a user account of a user of the user query. If the visual content is subsequently referenced in the dialog, the corresponding tokenized representation of the visual content is retrieved from the database. The corresponding tokenized representation of the visual content, corresponding tokenized representations of natural language content, and optionally other metadata can be processed, using a generative model, to generate a response responsive to the user input.
    Type: Grant
    Filed: June 26, 2024
    Date of Patent: June 9, 2026
    Assignee: GOOGLE LLC
    Inventors: Ágoston Weisz, Alessandro Agostini, François-Xavier Aubet, Khalid Salama, Trevor Strohman, Ilia Akolzin, Petre Petrov
  • Publication number: 20260155137
    Abstract: A method includes receiving a textual prompt directed toward a large language model (LLM)-powered assistant. The method also includes determining the textual prompt was generated by an automatic recognition system (ASR) system and, based on determining the textual prompt was generated by the ASR system, structuring a speech misrecognition awareness prompt. Here, the speech misrecognition awareness prompt includes: an awareness message that informs the LLM-powered assistant that the text prompt was generated by the ASR system and may be prone to speech recognition errors; and one or more error-correction pairs where each error-correction pair includes a corresponding misrecognized phrase and a corresponding correction phrase that corrects the corresponding misrecognized phrase. The method also includes processing, using the LLM-powered assistant, the textual prompt conditioned on the speech misrecognition awareness prompt to fulfill performance of the task specified by the natural language query.
    Type: Application
    Filed: December 4, 2024
    Publication date: June 4, 2026
    Applicant: Google LLC
    Inventors: Khalid Salama, Antonious Mamdouh Girgis Bebawy
  • Publication number: 20260050747
    Abstract: Implementations relate to receiving a free-form natural language input associated with a client device; processing, using a first generative model (GM), first GM input to generate corresponding first GM output; determining, based on the first GM output, an initial query that includes placeholder(s); retrieving placeholder data that includes, for the placeholder(s), a corresponding set of variables and a set of probability values corresponding to the set of variables; determining, based on the initial query, a final query; and providing the final query for processing by the first GM or a second GM. Determining the final query includes, for the placeholder(s): selecting, based on the corresponding set of variables and the set of probability values corresponding to the set of variables, a variable from the corresponding set of variables; and replacing the placeholder(s) with the selected variable.
    Type: Application
    Filed: August 13, 2024
    Publication date: February 19, 2026
    Inventor: Khalid Salama
  • Publication number: 20260004071
    Abstract: Implementations relate to handling visual content across a multi-turn dialog. A user input that includes natural language content and visual content is received during the dialog. If the visual content is being received for the first time in the dialog, the visual content is processed to generate a corresponding tokenized representation of the visual content. The corresponding tokenized representation can be cached in a database in association with the dialog, or in association with a user account of a user of the user query. If the visual content is subsequently referenced in the dialog, the corresponding tokenized representation of the visual content is retrieved from the database. The corresponding tokenized representation of the visual content, corresponding tokenized representations of natural language content, and optionally other metadata can be processed, using a generative model, to generate a response responsive to the user input.
    Type: Application
    Filed: June 26, 2024
    Publication date: January 1, 2026
    Inventors: Ágoston Weisz, Alessandro Agostini, François-Xavier Aubet, Khalid Salama, Trevor Strohman, Ilia Akolzin, Petre Petrov
  • Patent number: 12400083
    Abstract: A method includes obtaining a set of training queries that each specify a corresponding operation to perform and include a corresponding plurality of speech recognition hypotheses that each represent a corresponding candidate transcription of the training query, and a corresponding ground-truth transcription of the training query. For each training query, the method includes processing, using an encoder of a neural semantic parsing (NSP) model, the corresponding plurality of speech recognition hypotheses to generate a corresponding NSP embedding, processing, using a transcription decoder, the corresponding NSP embedding to generate a corresponding predicted transcription, and determining a corresponding first loss based on the corresponding predicted transcription and the corresponding ground-truth transcription.
    Type: Grant
    Filed: June 20, 2023
    Date of Patent: August 26, 2025
    Assignee: Google LLC
    Inventors: Khalid Salama, Ágoston Weisz
  • Publication number: 20250258861
    Abstract: Implementations disclosed herein are directed to at least responding to an input query comprising a natural language query and an image using a vision and language model (VLM). The input NL query and image are processed to generate sub-images (referred to herein as “tiles”) of the input image that are relevant to the NL query. The tiles are processed by one or more image analysis models, such as image search engines, to generate image facts that relate to the tiles, e.g., the contents of a tile, identities of objects/people in the tile, or the like. The VLM processes the image tiles, the NL query, and the image facts to generate a response to the input query. The response is rendered at a client device.
    Type: Application
    Filed: February 12, 2024
    Publication date: August 14, 2025
    Inventors: Khalid Salama, Arkadiusz Socala
  • Publication number: 20240427997
    Abstract: A method includes obtaining a set of training queries that each specify a corresponding operation to perform and include a corresponding plurality of speech recognition hypotheses that each represent a corresponding candidate transcription of the training query, and a corresponding ground-truth transcription of the training query. For each training query, the method includes processing, using an encoder of a neural semantic parsing (NSP) model, the corresponding plurality of speech recognition hypotheses to generate a corresponding NSP embedding, processing, using a transcription decoder, the corresponding NSP embedding to generate a corresponding predicted transcription, and determining a corresponding first loss based on the corresponding predicted transcription and the corresponding ground-truth transcription.
    Type: Application
    Filed: June 20, 2023
    Publication date: December 26, 2024
    Applicant: Google LLC
    Inventors: Khalid Salama, Ágoston Weisz