Patents by Inventor Haolin Jin

Haolin Jin has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Patent number: 12681830
    Abstract: The systems and methods disclosed herein enable the dynamic selection of one or more AI models to generate an output in response to an input. The system receives, from a computing device, an output generation request including an input for the generation of an output using one or more models from a plurality of models. The system generates expected values for a set of output attributes of the output generation request. For each particular model in the plurality of models, the system determines the capabilities of the particular model, and dynamically select a subset of models from the plurality of models. The system dynamically selects a subset of available system resources to process the input included in the output generation request. The system generates the output by processing the input included in the output generation request using the selected subset of available system resources.
    Type: Grant
    Filed: August 22, 2024
    Date of Patent: July 14, 2026
    Inventors: Sourabh Deb, Jason Engelbrecht, Zheyu Wang, Haolin Jin, Payal Jain, Tariq Husayn Maonah, Mariusz Saternus, Daniel Lewandowski, Biraj Krushna Rath, Stuart Murray, Philip Davies
  • Publication number: 20260195430
    Abstract: Systems, methods, and devices for facilitating computational resource access by artificial intelligence (AI) agents through token-based allocation and multi-agent workflow optimization. The system generates tokens corresponding to computational resources including processing power, memory, storage, and bandwidth. AI agents submit resource requests with priority tokens, creating queues ordered by priority token quantity. Higher-priority bids receive preferential positions. The system transfers resource tokens to agents based on queue order, enabling resource access through token exchange. The system can receive user prompts indicating computational objectives and determines multiple AI agentic approaches comprising AI model sequences. After evaluating approaches against operational policies, the system generates resource utilization and performance estimates, executing a preferred approach balancing efficiency with output quality.
    Type: Application
    Filed: March 2, 2026
    Publication date: July 9, 2026
    Inventors: Ganesh Prasad Bhat, James Myers, Payal Jain, Tariq Husayn Maonah, Mariusz Saternus, Daniel Lewandowski, Biraj Krushna Rath, Stuart Murray, Philip Davies, Sourabh Deb, Jason Engelbrecht, Zheyu Wang, Haolin Jin, Avi Levin, Nimrod Barak, Miriam Silver
  • Publication number: 20260154526
    Abstract: The systems and methods disclosed herein orchestrate task execution among autonomous (or semi-autonomous) AI agentic models (“agents”) using a gateway router that dynamically coordinates the agents based on prompt characteristics, user context, and/or real-time operational factors. Received inputs (e.g., prompts) are segmented into subcomponents (e.g., sub-queries), which are routed/mapped to candidate agents based on the output parameters of the subcomponent (e.g., performance thresholds, cost thresholds) and operational parameters (e.g., cost, performance metric values, user access restrictions, timing restrictions) of each agent. The gateway router maintains dynamic routing data structures for each agent that are continuously updated based on environmental stimuli (e.g., geo-political stimuli, sensor stimuli, agent stimuli). For example, the gateway router causes agents to dynamically switch between rule engines identified by the routing tables in response to detecting environmental stimuli.
    Type: Application
    Filed: January 23, 2026
    Publication date: June 4, 2026
    Applicant: Citibank, N.A.
    Inventors: James MYERS, Ganesh Prasad BHAT, Sourabh DEB, Jason ENGELBRECHT, Zheyu WANG, Haolin JIN
  • Patent number: 12608455
    Abstract: Systems, methods, and devices for facilitating computational resource access by artificial intelligence (AI) agents through token-based allocation and multi-agent workflow optimization. The system generates tokens corresponding to computational resources including processing power, memory, storage, and bandwidth. AI agents submit resource requests with priority tokens, creating queues ordered by priority token quantity. Higher-priority bids receive preferential positions. The system transfers resource tokens to agents based on queue order, enabling resource access through token exchange. The system can receive user prompts indicating computational objectives and determines multiple AI agentic approaches comprising AI model sequences. After evaluating approaches against operational policies, the system generates resource utilization and performance estimates, executing a preferred approach balancing efficiency with output quality.
    Type: Grant
    Filed: September 12, 2025
    Date of Patent: April 21, 2026
    Assignee: Citibank, N.A.
    Inventors: Ganesh Prasad Bhat, James Myers, Payal Jain, Tariq Husayn Maonah, Mariusz Saternus, Daniel Lewandowski, Biraj Krushna Rath, Stuart Murray, Philip Davies, Sourabh Deb, Jason Engelbrecht, Zheyu Wang, Haolin Jin, Avi Levin, Nimrod Barak, Miriam Silver
  • Patent number: 12602418
    Abstract: Systems, methods, and devices that relate to intelligent query decomposition and parallel routing for specialized model processing are disclosed. In one example aspect, the system receives a query from a user comprising a request relating to a particular domain. The system determines, using a decomposition model, a set of sub-queries based on semantic boundaries, syntactics, tasks, relationships, and rules relating to particular domains. The system inputs the set of sub-queries into a routing model to determine a set of specialized models. For each sub-query, the system routes the sub-query to a respective specialized model, generates an output, and assigns a confidence score. The system detects conflicts among outputs using a conflict detection model configured to identify discrepancies. The system generates an aggregated output by combining outputs according to a weighted aggregation algorithm prioritizing higher confidence scores and conflict resolution rules, then displays the aggregated output.
    Type: Grant
    Filed: August 25, 2025
    Date of Patent: April 14, 2026
    Inventors: Ganesh Prasad Bhat, James Myers, Zheyu Wang, Haolin Jin, Sourabh Deb, Jason Ryan Engelbrecht, Payal Jain, Tariq Husayn Maonah, Mariusz Saternus, Daniel Lewandowski, Biraj Krushna Rath, Stuart Murray, Philip Davies, Julisia Jackson, Chamindra Desilva, Shardul Malviya, Wayne Liao, Deepak Jain, Samantha Cory, Vishal Mysore, Ramkumar Ayyadurai
  • Patent number: 12596738
    Abstract: Systems for explainable large language model routing with immutable audit trails are disclosed. The system receives a query and determines its characteristics including complexity, domain, regulatory constraints, and performance requirements. It retrieves profiles for multiple LLMs from a model matrix containing performance attributes, resource consumption, and compliance parameters. The system selects a particular LLM by balancing resource consumption with performance requirements, evaluating regulatory compliance, ranking LLMs based on these factors, and prioritizing models with successful processing history. The system generates a human-readable explanation of the selection including decision factors, rationale, and alternatives considered. Finally, it records the selection and explanation in a tamper-evident, immutable audit trail data structure.
    Type: Grant
    Filed: August 28, 2025
    Date of Patent: April 7, 2026
    Assignee: Citibank, N.A.
    Inventors: Ganesh Prasad Bhat, Zheyu Wang, Haolin Jin, Sourabh Deb, Jason Ryan Engelbrecht, Payal Jain, Tariq Husayn Maonah, Mariusz Saternus, Daniel Lewandowski, Biraj Krushna Rath, Stuart Murray, Philip Davies, Julisia Jackson, Chamindra Desilva, Shardul Malviya, Wayne Liao, Deepak Jain, Samantha Cory, Vishal Mysore, Ramkumar Ayyadurai, James Myers
  • Publication number: 20260093791
    Abstract: Systems and methods for restructuring prompts in order to improve accuracy of outputs from models are disclosed herein. The system receives a user prompt indicating a request for data. The system generates a first and second output using a model, the first output generated based on the user prompt and the second output generated based on pseudocode. The system compares the first and second outputs to determine a match accuracy between the two outputs. If the two outputs sufficiently match, the system approves the user prompt. If the two outputs do not sufficiently match, the system initiates a prompt restructuring process, whereby the user prompt is restructured using pseudocode to improve the accuracy of the first output. The process is repeated iteratively until the restructured user prompt generates a first output that sufficiently matches the second output generated based on the pseudocode.
    Type: Application
    Filed: December 5, 2025
    Publication date: April 2, 2026
    Inventors: Nigil Satish Jeyashekar, Jason Engelbrecht, Zheyu Wang, Haolin Jin, Payal Jain, Tariq Husayn Maonah, Mariusz SATERNUS, Daniel Lewandowski, Biraj Krushna Rath, Stuart Murray, Philip Davies, Sourabh Deb
  • Patent number: 12591793
    Abstract: The systems and methods disclosed herein orchestrate task execution among autonomous (or semi-autonomous) AI agentic models (“agents”) responsive to a received query by using a hierarchical semantic fingerprinting framework to generate semantic-aware fingerprints for the query and agents. Queries and descriptions of agent capabilities are processed by a series of hierarchical levels using locality-sensitive hash (LSH) functions, where subsequent layers encode more complex semantic information and generate longer hash values. The hash values are aggregated into a semantic fingerprint. A bloom filter cascade uses a series of increasingly accurate hierarchical bloom filters to reject agent fingerprints that differ from the query fingerprint. The remaining agent fingerprints are compared bitwise to the query fingerprint, and those closest to the query fingerprint are selected to generate a routing path for the query. Responses from the selected agents are aggregated into an output that is responsive to the input.
    Type: Grant
    Filed: September 11, 2025
    Date of Patent: March 31, 2026
    Assignee: Citibank, N.A.
    Inventors: Ganesh Prasad Bhat, James Myers, Sourabh Deb, Jason Engelbrecht, Zheyu Wang, Haolin Jin
  • Patent number: 12536406
    Abstract: The systems and methods disclosed herein orchestrate task execution among autonomous (or semi-autonomous) AI agentic models (“agents”) using a gateway router that dynamically coordinates the agents based on prompt characteristics, user context, and/or real-time operational factors. Received inputs (e.g., prompts) are segmented into subcomponents (e.g., sub-queries), which are routed/mapped to candidate agents based on the output parameters of the subcomponent (e.g., performance thresholds, cost thresholds) and operational parameters (e.g., cost, performance metric values, user access restrictions, timing restrictions) of each agent. The gateway router maintains dynamic routing data structures for each agent that are continuously updated based on environmental stimuli (e.g., geo-political stimuli, sensor stimuli, agent stimuli). For example, the gateway router causes agents to dynamically switch between rule engines identified by the routing tables in response to detecting environmental stimuli.
    Type: Grant
    Filed: July 24, 2025
    Date of Patent: January 27, 2026
    Inventors: James Myers, Ganesh Prasad Bhat, Sourabh Deb, Jason Engelbrecht, Zheyu Wang, Haolin Jin
  • Patent number: 12524508
    Abstract: Systems and methods for restructuring prompts in order to improve accuracy of outputs from models are disclosed herein. The system receives a user prompt indicating a request for data. The system generates a first and second output using a model, the first output generated based on the user prompt and the second output generated based on pseudocode. The system compares the first and second outputs to determine a match accuracy between the two outputs. If the two outputs sufficiently match, the system approves the user prompt. If the two outputs do not sufficiently match, the system initiates a prompt restructuring process, whereby the user prompt is restructured using pseudocode to improve the accuracy of the first output. The process is repeated iteratively until the restructured user prompt generates a first output that sufficiently matches the second output generated based on the pseudocode.
    Type: Grant
    Filed: April 24, 2025
    Date of Patent: January 13, 2026
    Inventors: Nigil Satish Jeyashekar, Jason Engelbrecht, Zheyu Wang, Haolin Jin, Payal Jain, Tariq Husayn Maonah, Mariusz Saternus, Daniel Lewandowski, Biraj Krushna Rath, Stuart Murray, Philip Davies, Sourabh Deb
  • Publication number: 20260010458
    Abstract: Systems, methods, and devices for facilitating computational resource access by artificial intelligence (AI) agents through token-based allocation and multi-agent workflow optimization. The system generates tokens corresponding to computational resources including processing power, memory, storage, and bandwidth. AI agents submit resource requests with priority tokens, creating queues ordered by priority token quantity. Higher-priority bids receive preferential positions. The system transfers resource tokens to agents based on queue order, enabling resource access through token exchange. The system can receive user prompts indicating computational objectives and determines multiple AI agentic approaches comprising AI model sequences. After evaluating approaches against operational policies, the system generates resource utilization and performance estimates, executing a preferred approach balancing efficiency with output quality.
    Type: Application
    Filed: September 12, 2025
    Publication date: January 8, 2026
    Inventors: Ganesh Prasad Bhat, James Myers, Payal Jain, Tariq Husayn Maonah, Mariusz Saternus, Daniel Lewandowski, Biraj Krushna Rath, Stuart Murray, Philip Davies, Sourabh Deb, Jason Engelbrecht, Zheyu Wang, Haolin Jin, Avi Levin, Nimrod Barak, Miriam Silver
  • Publication number: 20260004162
    Abstract: The systems and methods disclosed herein orchestrate task execution among autonomous (or semi-autonomous) AI agentic models (“agents”) responsive to a received query by using a hierarchical semantic fingerprinting framework to generate semantic-aware fingerprints for the query and agents. Queries and descriptions of agent capabilities are processed by a series of hierarchical levels using locality-sensitive hash (LSH) functions, where subsequent layers encode more complex semantic information and generate longer hash values. The hash values are aggregated into a semantic fingerprint. A bloom filter cascade uses a series of increasingly accurate hierarchical bloom filters to reject agent fingerprints that differ from the query fingerprint. The remaining agent fingerprints are compared bitwise to the query fingerprint, and those closest to the query fingerprint are selected to generate a routing path for the query. Responses from the selected agents are aggregated into an output that is responsive to the input.
    Type: Application
    Filed: September 11, 2025
    Publication date: January 1, 2026
    Applicant: Citibank, N.A.
    Inventors: Ganesh Prasad BHAT, James MYERS, Sourabh DEB, Jason ENGELBRECHT, Zheyu WANG, Haolin JIN
  • Publication number: 20250390565
    Abstract: Systems, methods, and devices for facilitating computational resource access by artificial intelligence (AI) agents through token-based allocation and multi-agent workflow optimization. The system generates tokens corresponding to computational resources including processing power, memory, storage, and bandwidth. AI agents submit resource requests with priority tokens, creating queues ordered by priority token quantity. Higher-priority bids receive preferential positions. The system transfers resource tokens to agents based on queue order, enabling resource access through token exchange. The system can receive user prompts indicating computational objectives and determines multiple AI agentic approaches comprising AI model sequences. After evaluating approaches against operational policies, the system generates resource utilization and performance estimates, executing a preferred approach balancing efficiency with output quality.
    Type: Application
    Filed: September 12, 2025
    Publication date: December 25, 2025
    Inventors: Ganesh Prasad Bhat, James Myers, Payal Jain, Tariq Husayn Maonah, Mariusz Saternus, Daniel Lewandowski, Biraj Krushna Rath, Stuart Murray, Philip Davies, Sourabh Deb, Jason Engelbrecht, Zheyu Wang, Haolin Jin, Avi Levin, Nimrod Barak, Miriam Silver
  • Publication number: 20250384072
    Abstract: Systems for explainable large language model routing with immutable audit trails are disclosed. The system receives a query and determines its characteristics including complexity, domain, regulatory constraints, and performance requirements. It retrieves profiles for multiple LLMs from a model matrix containing performance attributes, resource consumption, and compliance parameters. The system selects a particular LLM by balancing resource consumption with performance requirements, evaluating regulatory compliance, ranking LLMs based on these factors, and prioritizing models with successful processing history. The system generates a human-readable explanation of the selection including decision factors, rationale, and alternatives considered. Finally, it records the selection and explanation in a tamper-evident, immutable audit trail data structure.
    Type: Application
    Filed: August 28, 2025
    Publication date: December 18, 2025
    Inventors: Ganesh Prasad Bhat, Zheyu Wang, Haolin Jin, Sourabh Deb, Jason Ryan Engelbrecht, Payal Jain, Tariq Husayn Maonah, Mariusz Saternus, Daniel Lewandowski, Biraj Krushna Rath, Stuart Murray, Philip Davies, Julisia Jackson, Chamindra DESILVA, Shardul MALVIYA, Wayne LIAO, Deepak JAIN, Samantha CORY, Vishal MYSORE, Ramkumar AYYADURAI, James MYERS
  • Publication number: 20250383970
    Abstract: The systems and methods disclosed herein orchestrate task execution among autonomous (or semi-autonomous) AI agentic models (“agents”) responsive to a received query by using a hierarchical model cascade to classify queries into agent domains. Queries are processed iteratively by a series of hierarchical levels containing one or more AI models, where each layer is more complex and imposes fewer resource constraints. Each level generates a classification and a confidence score pertaining to the classification. A dynamic bypass mechanism analyzes the classifications and confidence scores at each level to dynamically determine if one or more levels of the hierarchy can be bypassed while resulting in an accurate classification. The final classifications are matched to one or more agents that process the query. Responses from the candidate agents are aggregated into an output that is responsive to the input.
    Type: Application
    Filed: September 10, 2025
    Publication date: December 18, 2025
    Applicant: Citibank, N.A.
    Inventors: Ganesh Prasad BHAT, James MYERS, Sourabh DEB, Jason ENGELBRECHT, Zheyu WANG, Haolin JIN
  • Publication number: 20250378099
    Abstract: Systems, methods, and devices that relate to intelligent query decomposition and parallel routing for specialized model processing are disclosed. In one example aspect, the system receives a query from a user comprising a request relating to a particular domain. The system determines, using a decomposition model, a set of sub-queries based on semantic boundaries, syntactics, tasks, relationships, and rules relating to particular domains. The system inputs the set of sub-queries into a routing model to determine a set of specialized models. For each sub-query, the system routes the sub-query to a respective specialized model, generates an output, and assigns a confidence score. The system detects conflicts among outputs using a conflict detection model configured to identify discrepancies. The system generates an aggregated output by combining outputs according to a weighted aggregation algorithm prioritizing higher confidence scores and conflict resolution rules, then displays the aggregated output.
    Type: Application
    Filed: August 25, 2025
    Publication date: December 11, 2025
    Inventors: Ganesh Prasad Bhat, James Myers, Zheyu Wang, Haolin Jin, Sourabh Deb, Jason Ryan Engelbrecht, Payal Jain, Tariq Husayn Maonah, Mariusz Saternus, Daniel Lewandowski, Biraj Krushna Rath, Stuart Murray, Philip Davies, Julisia Jackson, Chamindra DESILVA, Shardul MALVIYA, Wayne LIAO, Deepak JAIN, Samantha CORY, Vishal MYSORE, Ramkumar AYYADURAI
  • Publication number: 20250371433
    Abstract: Systems, methods, and devices that relate to routing requests to large language models (LLMs) are disclosed. In one example aspect, the system receives session-specific data elements in response to a request to generate an output using LLMs. The system determines a hierarchy of operational constraints including privacy protocols and performance requirements. Weights for a multi-variable optimization are dynamically updated using the session-specific data elements. The system executes the multi-variable optimization across candidate LLMs that satisfy privacy constraints and optimize performance constraints. Based on the optimization, at least one candidate LLM is selected and the request is routed to it. In response to performance feedback, the system automatically selects a different LLM to improve one constraint, resulting in degradation of another constraint.
    Type: Application
    Filed: August 15, 2025
    Publication date: December 4, 2025
    Inventors: Ganesh Prasad Bhat, Zheyu Wang, Haolin Jin, Sourabh Deb, Jason Ryan Engelbrecht, Payal Jain, Tariq Husayn Maonah, Mariusz Saternus, Daniel Lewandowski, Biraj Krushna Rath, Stuart Murray, Philip Davies, James Myers
  • Publication number: 20250348707
    Abstract: The systems and methods disclosed herein orchestrate task execution among autonomous (or semi-autonomous) AI agentic models (“agents”) using a gateway router that dynamically coordinates the agents based on prompt characteristics, user context, and/or real-time operational factors. Received inputs (e.g., prompts) are segmented into subcomponents (e.g., sub-queries), which are routed/mapped to candidate agents based on the output parameters of the subcomponent (e.g., performance thresholds, cost thresholds) and operational parameters (e.g., cost, performance metric values, user access restrictions, timing restrictions) of each agent. The gateway router maintains dynamic routing data structures for each agent that are continuously updated based on environmental stimuli (e.g., geo-political stimuli, sensor stimuli, agent stimuli). For example, the gateway router causes agents to dynamically switch between rule engines identified by the routing tables in response to detecting environmental stimuli.
    Type: Application
    Filed: July 24, 2025
    Publication date: November 13, 2025
    Applicant: Citibank, N.A.
    Inventors: James MYERS, Ganesh Prasad BHAT, Sourabh DEB, Jason ENGELBRECHT, Zheyu WANG, Haolin JIN
  • Publication number: 20250321850
    Abstract: The systems and methods disclosed herein enable the dynamic selection of one or more AI models to generate an output in response to an input. The system receives, from a computing device, an output generation request including an input for the generation of an output using one or more models from a plurality of models. The system generates expected values for a set of output attributes of the output generation request. For each particular model in the plurality of models, the system determines the capabilities of the particular model, and dynamically select a subset of models from the plurality of models. The system dynamically selects a subset of available system resources to process the input included in the output generation request. The system generates the output by processing the input included in the output generation request using the selected subset of available system resources.
    Type: Application
    Filed: August 22, 2024
    Publication date: October 16, 2025
    Inventors: Sourabh Deb, Jason Engelbrecht, Zheyu Wang, Haolin Jin
  • Publication number: 20250322047
    Abstract: Systems and methods for restructuring prompts in order to improve accuracy of outputs from models are disclosed herein. The system receives a user prompt indicating a request for data. The system generates a first and second output using a model, the first output generated based on the user prompt and the second output generated based on pseudocode. The system compares the first and second outputs to determine a match accuracy between the two outputs. If the two outputs sufficiently match, the system approves the user prompt. If the two outputs do not sufficiently match, the system initiates a prompt restructuring process, whereby the user prompt is restructured using pseudocode to improve the accuracy of the first output. The process is repeated iteratively until the restructured user prompt generates a first output that sufficiently matches the second output generated based on the pseudocode.
    Type: Application
    Filed: April 24, 2025
    Publication date: October 16, 2025
    Inventors: Nigil Satish Jeyashekar, Jason Engelbrecht, Zheyu Wang, Haolin Jin, Payal Jain, Tariq Husayn Maonah, Mariusz SATERNUS, Daniel Lewandowski, Biraj Krushna Rath, Stuart Murray, Philip Davies, Sourabh Deb