Patents by Inventor Wendi Cui

Wendi Cui has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20260187196
    Abstract: Aspects of the present disclosure relate to taxonomy guided retrieval augmented generation in machine learning models. Embodiments include generating a summary of each respective document in a plurality of documents and extracting attributes and relationships between the attributes from each respective summary. Embodiments include generating a hierarchy of taxonomies using a first graph and associating each document of the plurality of documents with a corresponding taxonomy of the taxonomies based on the hierarchy. Embodiments include creating an embedding associated with each document based on contents of each document and the associating and constructing a second graph containing the embedding associated with each document. Embodiments include retrieving one or more inputs to provide to a machine learning model in connection with a query based on comparing an embedding of the query to corresponding embeddings in the graph and generating a response to the query based on the one or more inputs.
    Type: Application
    Filed: December 27, 2024
    Publication date: July 2, 2026
    Inventors: Wendi CUI, Jiaxin ZHANG, Damien J. LOPEZ, Colin P. RYAN
  • Publication number: 20260119537
    Abstract: Certain aspects of the disclosure provide a method for augmenting an information repository for information retrieval by a large language model. The method may include obtaining an information item from an information repository; allocating portions of the information item into a plurality of contextual units; generating, by a first language model, a plurality of content units based on the plurality of contextual units; constructing an augmented dataset that includes one or more content units of the plurality of content units; receiving a query; inputting the query and the one or more content units of the plurality of content units into one or more language models; and obtaining, as output from the one or more language models, a response to the query based on the input query and the one or more content units of the plurality of content units.
    Type: Application
    Filed: October 29, 2024
    Publication date: April 30, 2026
    Inventors: Wendi CUI, Jiaxin ZHANG, Yiran HUANG, Damien LOPEZ
  • Publication number: 20260119803
    Abstract: Certain embodiments of the disclosure provide techniques for knowledge refinement for language model fine-tuning. A method generally includes obtaining a raw information item; partitioning the raw information item into a plurality of first contextual units, wherein each first contextual unit comprises a first portion of the raw information item; generating, via a first language model, first synthetic data based on: the plurality of first contextual units; the raw information item; and at least one of fine-grained synthesis, interleaved generation, or assembly augmentation; and fine-tuning a second language model based on the first synthetic data.
    Type: Application
    Filed: October 25, 2024
    Publication date: April 30, 2026
    Inventors: Jiaxin ZHANG, Wendi CUI, Yiran HUANG, Kamalika DAS, Sricharan KALLUR PALLI KUMAR
  • Patent number: 12585676
    Abstract: A large language model (LLM) is refined using an unsupervised alignment process for enhanced reasoning and deep thinking capabilities. An iterative search and optimization procedure is used to explore the space of possible thought generations enabling the model to improve deep thinking capabilities without direct supervision. The LLM is prompted to produce a first thought and response pair in response to a query. The thought and response pair is evaluated with a trained judge LLM, which provides feedback for the thought and for the response. A revised prompt is generated in response to the feedback and provided to the LLM, which produces second thought and response pair. The LLM is refined based on the first, non-preferred, thought and response pair and the second, preferred, thought and response pair to align with the preferred thought and response pair, e.g., using a direct preference optimization process.
    Type: Grant
    Filed: May 30, 2025
    Date of Patent: March 24, 2026
    Assignee: Intuit Inc.
    Inventors: Ankita Sinha, Jiaxin Zhang, Kamalika Das, Sricharan Kallur Palli Kumar, Wendi Cui
  • Publication number: 20260064745
    Abstract: A method includes receiving a query from a user device and applying an embedding model to the query to generate a query vector data structure. A vector comparator is applied to the query vector data structure and embedded document chunk vector data structures to output a score measuring a semantic similarity between the vector data structures. The embedded document chunk vector data structures are generated by the embedding model being applied to documents in a knowledge base corpus. Responsive to the score failing to satisfy a threshold value, the query is rejected as being an out-of-distribution query by transmitting an electronic reject message. Responsive to the score satisfying the threshold value and indicating that the query is an in-distribution query, the language model is applied to the query and to the knowledge base corpus to output an answer. The method also includes returning the answer to the user device.
    Type: Application
    Filed: August 28, 2024
    Publication date: March 5, 2026
    Applicant: Intuit Inc.
    Inventors: Jiaxin ZHANG, Wendi CUI, Kamalika DAS, Sricharan KALLUR PALLI KUMAR
  • Patent number: 12536449
    Abstract: Certain aspects of the disclosure provide a method for updating a retrieval augmented generation (RAG) system. The method includes receiving a user query and retrieving a set of datasets from an external knowledge base associated with a language model. The user query and retrieved datasets are provided to the language model as input tokens, which generates a response comprising output tokens. The system then extracts cross-attention weights from the language model, indicating how much attention each output token paid to each input token. Using these weights, the system generates attention scores for each dataset and identifies a top-k set of most attended datasets. If the generated response is determined to be relevant to the user query, the top-k most attended datasets are labeled as positive examples. The system then updates its parameters to prioritize retrieving these positive examples for future queries, enabling continuous self-supervised improvement.
    Type: Grant
    Filed: August 28, 2025
    Date of Patent: January 27, 2026
    Assignee: Intuit Inc.
    Inventors: Wendi Cui, Damien J. Lopez, Amit Agarwal, Joel D. Atwood, Palak Samel
  • Publication number: 20260023763
    Abstract: A method includes performing a semantic mutation on an initial prompt by a prompt optimizer large language model (LLM) to obtain an initial generation of prompts. The method further includes evaluating the initial generation of prompts using a pareto selection function, to obtain a first generation of prompts. The method further includes mutating the first generation of prompts according to a first objective. The method further includes mutating a second generation of prompts obtained from the first generation of prompts according to a second objective. The method further includes performing a crossover mutation on a generation of parent prompts obtained from the second generation of prompts to obtain a result population of prompts. The method further includes adding the result population of prompts to a prompt population.
    Type: Application
    Filed: July 18, 2024
    Publication date: January 22, 2026
    Applicant: Intuit Inc.
    Inventors: Ankita SINHA, Jiaxin ZHANG, Wendi CUI, Kamalika DAS
  • Publication number: 20260004158
    Abstract: Certain aspects of the disclosure provide unified in-context prompt optimization for large language models that achieves joint optimization of prompt instruction and examples. A multi-phase approach is provided that includes multiple mutation operations. Further, the approach alternates between optimization strategies for exploration for global search and exploitation for local search. Global initialization creates a diverse set of candidate prompts based on the availability of data and utilizing Lamarckian or semantic mutation. Local feedback mutation, global evolution mutation, and local semantic mutation can subsequently be employed iteratively to generate a revised set of candidate prompts. A prompt from the revised set of candidate prompts can be selected based on an evaluation of the candidate prompts. Subsequently, the selected prompt can be output for a machine-learning task.
    Type: Application
    Filed: June 28, 2024
    Publication date: January 1, 2026
    Inventors: Wendi CUI, Jiaxin ZHANG, Damien LOPEZ, Kamalika DAS, Sricharan Kallur Palli KUMAR
  • Publication number: 20250371279
    Abstract: A method further includes performing the following phases for at least two iterations. In a first phase, the method includes, iteratively, applying an evolutionary algorithm to a current instruction to generate a revised instruction, testing, by applying a large language model (LLM), prompts including the current instruction and the revised instruction, respectively, training examples selected by the example selector to obtain test results, comparing the test results to obtain a comparison result, setting the revised instruction as the current instruction, and exiting the first phase when the comparison result satisfies a first phase stop condition. In a second phase, the method further includes selecting, by the example selector, training examples, testing, using the current instruction, the training examples to obtain a test result, and modifying, after executing the first phase, the example selector based on the test result.
    Type: Application
    Filed: May 30, 2024
    Publication date: December 4, 2025
    Applicant: Intuit Inc.
    Inventors: Jiaxin ZHANG, Xiang GAO, Wendi CUI, Na XU, Maya Vered LIVSHITS, Sourav PROSAD, Arkadeep BANERJEE, Kun HU, Andrew MATTARELLA-MICKE, Vignesh Thirukazhukundram SUBRAHMANIAM, Kamalika DAS
  • Patent number: 12468747
    Abstract: This invention relates to systems and methods for performing double-level ranking of documents. The system implements methods for retrieving documents based on a pre-processed user query to generate a document level ranking of one or more documents that are determined to be relevant. The system implements methods for aggregating one or more sub-topic snippets from the document level ranked documents. The system further implements methods for generating a topic level ranking of the one or more sub-topic snippets. Once topic level ranking of the one or more sub-topic snippets has been performed, the system implements methods for transmitting the topic level ranked one or more sub-topic snippets to a user associated with the user query.
    Type: Grant
    Filed: February 28, 2024
    Date of Patent: November 11, 2025
    Assignee: INTUIT INC.
    Inventors: Wendi Cui, Colin P. Ryan, Damien J. Lopez
  • Publication number: 20250335773
    Abstract: A method includes performing a gradient descent mutation of a current generation of prompts by an evolutionary algorithm framework engine. The gradient descent mutation includes sending a prompt to a large language model (LLM) with an evaluation input-output pair and instructing the LLM to generate a modification recommendation for the prompt. The prompt is modified according to the modification recommendation. The modified prompt is processed by the LLM with the evaluation input output pair, causing the LLM to generate a response matching the output of the evaluation input-output pair. The modified prompt is added to a next generation of prompts.
    Type: Application
    Filed: April 30, 2024
    Publication date: October 30, 2025
    Applicant: Intuit Inc.
    Inventors: Wendi CUI, Jiaxin ZHANG, Kamalika DAS, Damien J. LOPEZ, Sricharan Kallur Palli KUMAR
  • Publication number: 20250322243
    Abstract: A system and method for optimizing instructions for Large Language Models (LLMs). The system and method comprise a text embedding model configured to convert the text prompt into an embedding space. An optimization module is configured to optimize the prompt in the embedding space, and an invertible embedding model is configured to decode the optimized embedding space back into a text prompt that may be input to a blackbox LLM.
    Type: Application
    Filed: April 11, 2024
    Publication date: October 16, 2025
    Applicant: INTUIT INC.
    Inventors: Jiaxin ZHANG, Wendi CUI, Kamalika DAS, Sricharan Kallur Palli KUMAR
  • Patent number: 12423313
    Abstract: A method includes obtaining a hierarchical document structure of a raw document. The hierarchical document structure includes a multitude of sections arranged in a hierarchy of successive document levels. A hierarchical document graph having a graph hierarchical structure corresponding to the hierarchical document structure of the raw document is constructed. A hierarchical document graph corresponding to the raw document matching a user query is retrieved. A hierarchical search is performed on the hierarchical document graph to obtain a set of relevant nodes. A set of relevant content embeddings is retrieved from the set of relevant nodes. A large language model (LLM) generates a response to the user query from the set of relevant content embeddings. The response is presented in a user application.
    Type: Grant
    Filed: April 30, 2025
    Date of Patent: September 23, 2025
    Assignee: Intuit Inc.
    Inventors: Wendi Cui, Jiaxin Zhang, Qi Shen, Beatrice Hendra
  • Publication number: 20250272321
    Abstract: This invention relates to systems and methods for performing double-level ranking of documents. The system implements methods for retrieving documents based on a pre-processed user query to generate a document level ranking of one or more documents that are determined to be relevant. The system implements methods for aggregating one or more sub-topic snippets from the document level ranked documents. The system further implements methods for generating a topic level ranking of the one or more sub-topic snippets. Once topic level ranking of the one or more sub-topic snippets has been performed, the system implements methods for transmitting the topic level ranked one or more sub-topic snippets to a user associated with the user query.
    Type: Application
    Filed: February 28, 2024
    Publication date: August 28, 2025
    Applicant: INTUIT INC.
    Inventors: Wendi CUI, Colin P. Ryan, Damien J. LOPEZ
  • Patent number: 12361070
    Abstract: A method includes receiving a query and applying a classifier to a hierarchical taxonomy and the query to output a query taxonomy array. Each of a number of taxonomy tags has an associated level in a hierarchical taxonomy. The query taxonomy array includes query tags selected, based on a term contained in the query, from the taxonomy tags. The query taxonomy array is compared to article tags associated with articles. The article tags are selected from the taxonomy tags. Shared tags are also identified based on sharing. A list of articles is generated, the list including a subset of articles selected from the articles. The subset of articles are associated with the shared tags. Corresponding scores are assigned to the subset of articles in the list. The list is sorted, based on the corresponding scores, to generate a sorted list of articles. The sorted list is presented.
    Type: Grant
    Filed: June 24, 2024
    Date of Patent: July 15, 2025
    Assignee: Intuit Inc.
    Inventors: Wendi Cui, Damien J. Lopez, Colin P. Ryan
  • Publication number: 20250139373
    Abstract: Output sentences of a primary large language model is provided to a criteria model including a second large language model. The criteria model compares the output to a reference source. As a result of comparing, the criteria model generates a first data structure including a first vector. The first vector stores, an evaluation of the output as being consistent or inconsistent with the reference source, and a corresponding reason for the evaluation. The criteria model identifies an inconsistent sentence, in the sentences, that is inconsistent with the reference source. The method also includes rewriting, by a reason improver model including a third large language model, the inconsistent sentence into a consistent sentence. The consistent sentence is consistent with the reference source. The output is modified by replacing the inconsistent sentence in the sentences with the consistent sentence. Modifying generates a modified output. The method also includes returning the modified output.
    Type: Application
    Filed: October 31, 2024
    Publication date: May 1, 2025
    Applicant: Intuit Inc.
    Inventors: Wendi CUI, Jiaxin ZHANG, Damien LOPEZ, Kamalika DAS, Sricharan KUMAR
  • Publication number: 20250139374
    Abstract: Providing an output of a primary large language model to a criteria model including a second large language model. The criteria model compares each of the sentences to a reference source and generates a first data structure including a first vector. The first vector stores, for each of the sentences, a corresponding evaluation of a given sentence as being consistent or inconsistent with the reference source, and a corresponding reason for the corresponding evaluation of the given sentence. The first data structure is provided to a converter model including a third large language model. The converter model converts the first data structure to a second data structure. The second data structure includes a second vector storing scores indicating a corresponding consistency value for each of the sentences. A metric, indicating an overall consistency of the output with respect to the reference source, is generated from the second data structure.
    Type: Application
    Filed: October 31, 2024
    Publication date: May 1, 2025
    Applicant: Intuit Inc.
    Inventors: Wendi CUI, Jiaxin ZHANG, Damien LOPEZ, Colin RYAN
  • Publication number: 20250077940
    Abstract: Certain aspects of the disclosure provide systems and methods for detecting hallucinations in machine learning models. A method generally includes generating a potential answer from an initial prompt received from a user. The method generally includes interrogating the machine learning model with a verification prompt formulated to elicit a positive or negative response from the machine learning model based on the potential answer and initial prompt. A negative response by the neural network model to the verification prompt is indicative of the potential answer being a hallucination. A positive response by the neural network model to the verification prompt is indicative of the potential answer being free from a hallucination. The method generally includes outputting to the user the potential answer as a final answer upon receiving a positive response to the verification prompt.
    Type: Application
    Filed: August 31, 2023
    Publication date: March 6, 2025
    Inventors: Wendi CUI, Colin P. RYAN, Damien J. LOPEZ, Palak SAMEL, Joel D ATWOOD
  • Patent number: 12147485
    Abstract: Certain aspects of the present disclosure provide techniques for managing a search engine based on search performance metrics. An example method generally includes dividing a set of search history data into a first subset of search history data and a second subset of search history data. The first subset of data is associated with interaction with search results, and the second subset of data is associated with non-interaction with search results. A first quality score is generated for searches in the first subset of data. A second quality score is generated for searches in the second subset of data based on different search intents identified for each temporally related group in the second subset of data. An overall quality score is generated for a search engine, and one or more actions with respect to the search engine are taken based on the overall quality score.
    Type: Grant
    Filed: January 11, 2024
    Date of Patent: November 19, 2024
    Assignee: INTUIT INC.
    Inventors: Wendi Cui, Damien J. Lopez, Colin P. Ryan
  • Publication number: 20240143680
    Abstract: Certain aspects of the present disclosure provide techniques for managing a search engine based on search performance metrics. An example method generally includes dividing a set of search history data into a first subset of search history data and a second subset of search history data. The first subset of data is associated with interaction with search results, and the second subset of data is associated with non-interaction with search results. A first quality score is generated for searches in the first subset of data. A second quality score is generated for searches in the second subset of data based on different search intents identified for each temporally related group in the second subset of data. An overall quality score is generated for a search engine, and one or more actions with respect to the search engine are taken based on the overall quality score.
    Type: Application
    Filed: January 11, 2024
    Publication date: May 2, 2024
    Inventors: Wendi CUI, Damien J. LOPEZ, Colin P. RYAN