Patents Examined by Mehran Kamran
-
Patent number: 12717555Abstract: A system for large language model (LLM) deployment dynamically provisions GPU resources across availability zones (AZs) to ensure scalable and cost-effective inference. The system initially provisions a node in a first cluster within a first AZ and receives a request to deploy an LLM. The system determines the GPU resource requirements for operating the LLM, then collects and analyzes real-time metrics—such as availability, pricing, and performance—across multiple alternative AZs containing diverse compute resources. Based on this analysis, the system selects a second AZ optimized for GPU resource efficiency and provisions a second node in a separate cluster. This second node is logically registered as a virtual node within the original cluster, allowing seamless deployment of the LLM while maintaining transparency and performance.Type: GrantFiled: May 7, 2025Date of Patent: August 25, 2026Assignee: CAST AI Group, Inc.Inventor: Leonid Kuperman
-
Patent number: 12717615Abstract: Systems and methods for processing a distributed transaction are provided. A distributed execution statement including a plurality of tasks for executing a query and a table to be modified by the executed query is received. A first task is transmitted to a first backend node and a second task is transmitted to a second backend node. A first confirmation that the first task is executed and a second confirmation that the second task is executed are received. Executing the first task includes making a first modification to the table and a first identification of the first modification, and executing the second task includes making a second modification to the table and a second identification of the second modification. Based at least on each of the first task and the second task being completed, committing the first modification and the second modification to the table.Type: GrantFiled: June 20, 2023Date of Patent: August 25, 2026Assignee: Microsoft Technology Licensing, LLC.Inventors: Alan Dale Halverson, Ryan Daniel O'Connor, Jose Aguilar Saborit, Raghunath Ramakrishnan
-
Patent number: 12710983Abstract: A method for allocating computing resources for a vehicle includes determining an optimal task configuration for a computing task based at least in part on a task constraint of the computing task. The method further may include determining a criticality level of the computing task based at least in part on the optimal task configuration and the task constraint of the computing task. The method further may include routing the computing task to one of a plurality of remote server systems based at least in part on the criticality level of the computing task.Type: GrantFiled: September 6, 2023Date of Patent: August 18, 2026Assignee: GM GLOBAL TECHNOLOGY OPERATIONS LLCInventors: Angelos Angelopoulos, Arun Adiththan, Md Mhafuzul Islam, Paolo Giusto, Adi Enzel
-
Patent number: 12670016Abstract: Methods and apparatus for hardware support for low latency microservice deployments in switches. A switch is communicatively coupled via a network or fabric to a plurality of platforms configured to implement one or more microservices. The microservices are used to perform a distributed workload, job, or task as defined by a corresponding graph representation of the microservices including vertices (also referred to as nodes) associated with microservices and edges defining communication between microservices. The graph representation also defines dependencies between microservices. The switch is configured to schedule execution of the graph of microservices on the plurality of platforms, including generating an initial schedule that is dynamically revised during runtime in consideration of performance telemetry data for the microservices received from the platforms and network/fabric utilization monitored onboard the switch.Type: GrantFiled: March 21, 2022Date of Patent: June 30, 2026Assignee: Intel CorporationInventors: Francesc Guim Bernat, Karthik Kumar, Alexander Bachmutsky
-
Patent number: 12664027Abstract: Aspects of the disclosure relate to real-time management of data container generation, authorization, and throttling. In some embodiments, a computing platform may receive and store data container management files, each associated with a different data container. The computing platform may thereafter receive a first data container and retrieve a first data container management file associated with the first data container. The computing platform may perform a preliminary analysis of the first data container to determine whether the first data container is complete. If the first data container is incomplete, the computing platform may send one or more analysis messages to a data container generation module to queue the first data container and retrieve an identification element associated with a missing data set. The computing platform may then retrieve the missing data set and generate an updated first data container by supplementing the first data container with the missing data set.Type: GrantFiled: March 23, 2023Date of Patent: June 23, 2026Assignee: Bank of America CorporationInventors: Manu Kurian, Paul Roscoe, Padmanabhan Iyer, Mahesh Bhashetty
-
Patent number: 12664012Abstract: Embodiments herein relate to providing uniform servicing of workloads at a set of servers in a computer network. A platform determines and meets the performance requirements of a workload by scaling a performance capability of a group of processing units such as central processing units (CPUs) which are assigned to service the workload. This can involve increasing the power (P) state of one or more of the processing units to a highest P state in the group, so that every processing units in the group provides the same performance for a given workload. The platform can manage scaling of the processing units performance by reading a performance profile list which indicates minimum and maximum scaling points for programs that are executed to service the workload.Type: GrantFiled: September 30, 2022Date of Patent: June 23, 2026Assignee: Intel CorporationInventors: Subhankar Panda, Rupal M. Parikh, Gaurav Porwal, Raghavendra Nagaraj, Sagar C. Pawar, Prakash Pillai
-
Patent number: 12664028Abstract: Example capacity adjustment methods, computing devices, and computer storage media are provided. One example capacity adjustment method includes creating a scaling group. A first instance is created on a first server set for the scaling group. A second instance is created on a second server set for the scaling group. A quantity of instances deployed on the first server set for the scaling group is limited by an upper limit value.Type: GrantFiled: January 6, 2023Date of Patent: June 23, 2026Assignee: Huawei Technologies Co., Ltd.Inventors: Guangcheng Li, Xi Chen
-
System and method for managing resource elasticity for a production environment with logical devices
Patent number: 12664029Abstract: A method for managing a production environment includes obtaining a notification for an alert event for the production environment, in response to the notification: making a determination that a resource limitation is exceeded by the alert event, based on the determination, performing a resource elasticity process based on a priority list to increase the resource limitation to a new resource limitation such that the alert event does not exceed the new resource limitation, wherein the resource limitation is based on available computing resources in the production environment, and resolving the alert event after the increase of the resource limitation.Type: GrantFiled: January 18, 2023Date of Patent: June 23, 2026Assignee: Dell Products L.P.Inventor: Parminder Singh Sethi -
Patent number: 12645477Abstract: The disclosure provides an approach for cross-cluster service resource discovery. A method includes obtaining, at a common store in a first node cluster in a cluster set information about a service resource of a second node cluster. The method includes creating a multi-cluster object associated with the service resource, wherein the multi-cluster object provides an association between the service resource and one or more endpoints on the second node cluster. The method includes storing the multi-cluster object in the common store, wherein the multi-cluster object is accessible in the common store by any of the plurality of node clusters in the cluster set to access the service resource on any of the one or more endpoints on the second node cluster.Type: GrantFiled: July 28, 2022Date of Patent: June 2, 2026Assignee: VMware, Inc.Inventors: Lan Luo, Wenfeng Liu, Donghai Han, Jianjun Shen
-
Patent number: 12639090Abstract: The present disclosure is a new and innovative system, methods and apparatus live storage migration. In an example, a system includes a memory and processor in communication with the memory. The processor is configured to receive a request to perform live storage migration of a guest managed by a source hypervisor on a source machine to a destination machine. The guest is configured to store data in blocks of block storage. The source hypervisor, executing on the processor, receives a hint for each block of data in block storage of the guest via an agent of the guest, wherein the hint relates to properties of a specific block of data. The source hypervisor then determines an efficient prioritization that identifies which blocks to copy and in what order as a part of the live storage migration based on the hints received from the agent. The source hypervisor then copies the blocks of data identified for copying in the migration based on the prioritization, saving time and computational resources.Type: GrantFiled: December 1, 2022Date of Patent: May 26, 2026Assignee: Red Hat, Inc.Inventor: Yaniv Kaul
-
Patent number: 12613736Abstract: A system for executing a plurality of software threads, comprising: a plurality of processing circuitries; a plurality of memory areas connected to the processing circuitries, each memory area associated with at least one of the processing circuitries; and at least one hardware processor, connected to the processing circuitries and configured for: in each of a plurality of iterations: while the processing circuitries execute the software threads, collecting for each thread a plurality of thread statistical values indicative of a plurality of memory accesses to at least some of the memory areas performed when executing the thread; for at least one thread, performing an analysis comprising the thread statistical values thereof to identify a preferred memory area of the plurality of memory areas; and configuring one of the at least one processing circuitry associated with the preferred memory area to execute the at least one thread.Type: GrantFiled: July 25, 2022Date of Patent: April 28, 2026Assignee: Next Silicon LtdInventors: Elad Raz, Ilan Tayari
-
Patent number: 12613857Abstract: Aspects of the present disclosure are directed to a system comprising a memory having computer-readable instructions stored thereon, and a processor of a database server, the processor executing the computer-readable instructions to generate a request to a control plane for an operation to be performed on the database server, wherein the control plane is configured to communicate with a plurality of database servers having a plurality of agents running thereon, and wherein each of the plurality of agents has a dedicated communication connection with the control plane, publish the request on the dedicated communication connection associated with the agent to send the request to the control plane, receive, on the dedicated communication connection, a response from the control plane, the response comprising a response to the request from a service of the control plane, and execute the operation on the database server based on the response.Type: GrantFiled: May 24, 2023Date of Patent: April 28, 2026Assignee: Nutanix, Inc.Inventors: Nilesh Vaishnav, Shurya Kumar N S, Akshay Chandak, Vaibhaw Pandey
-
Patent number: 12591444Abstract: A system controls access to a physical address (PA) space. The system includes multiple system resources addressable within the PA space, and multiple processing circuits executing multiple virtual machines (VMs). A given region of the PA space is dedicated to addressing the VMs. The system also includes multiple memory management units (MMUs) coupled to corresponding processing circuits. A given MMU is operative to translate a virtual address indicated in an access request from a processing circuit into a requested PA that is accessible by the processing circuit according to a configurable setting of the given MMU. The system further includes multiple memory protection units (MPUs). A given MPU, which is coupled to a target system resource allocated with the requested PA, is operative to grant or deny the request based on information indicating whether the requested PA is accessible to a requesting VM executed on the processing circuit.Type: GrantFiled: August 8, 2022Date of Patent: March 31, 2026Assignee: MediaTek Inc.Inventors: Chih-Hsiang Hsiao, Hung-Wen Chien
-
Patent number: 12585486Abstract: In some implementations, an orchestration system may receive information regarding a containerized network function (CNF) to be deployed. The orchestration system may obtain deployment information regarding requirements for deploying the CNF, wherein the requirements include one or more of a computing requirement, a memory requirement, or a network connectivity requirement. The orchestration system may provide a request to obtain one or more recommendations regarding one or more schedulers configured to deploy CNFs associated with the requirements. The orchestration system may obtain the one or more recommendations regarding the one or more schedulers based on providing the request.Type: GrantFiled: December 16, 2022Date of Patent: March 24, 2026Assignee: Verizon Patent and Licensing Inc.Inventors: Abhishek Kumar, Hans Raj Nahata, Neeraj Bhatt
-
Patent number: 12585497Abstract: A device may queue a thread for execution by a thread pool, including configuring an ambient context with a cancellation token. Based on a call of a synchronous method by the thread, the device may determine that the thread uses a green threads model and identify the cancellation token from the ambient context. Based on the thread using the green threads model, the device may call an asynchronous method that uses an asynchronous operation to perform a task of the synchronous method, including passing the asynchronous method the cancellation token. Based on identifying a state change of the cancellation token, the device may terminate the asynchronous operation.Type: GrantFiled: May 1, 2023Date of Patent: March 24, 2026Assignee: Microsoft Technology Licensing, LLCInventors: Stephen Harris Toub, David Charles Wrighton
-
Patent number: 12572394Abstract: In a Boundaryless Control High Availability (“BCHA”) system (e.g., industrial control system) comprising multiple computing resources (or computational engines) running on multiple machines, technology for computing in real time the overall system availability based upon the capabilities/characteristics of the available computing resources, applications to execute and the distribution of the applications across those resources is disclosed. In some embodiments, the disclosed technology can dynamically manage, coordinate recommend certain actions to system operators to maintain availability of the overall system at a desired level. High Availability features may be implemented across a variety of different computing resources distributed across various aspects of a BCHA system and/or computing resources. Two example implementations of BCHA systems described involve an M:N working configuration and M:N+R working configuration.Type: GrantFiled: November 3, 2023Date of Patent: March 10, 2026Assignee: Schneider Electric Systems USA, Inc.Inventors: Raja Ramana Macha, Andrew Lee David Kling, Frans Middeldorp, Nestor Jesus Camino, Jr., James Gerard Luth, James P. Mcintyre
-
Patent number: 12561158Abstract: A method and system of deployment of a virtualized service on a cloud infrastructure are described. A first service function specification of a first service function is selected. A determination of a set of the computing systems and a set of the links is performed based on availability and characteristics of the computing systems and the network resources in the cloud infrastructure. A selection of a first computing system to be assigned to host the first service function and links is performed based on the first service function specification and based on interoperability requirements for the first service function and one or more other ones of the service functions that form the virtualized service. The selection of a service function and the determination of a computing system and links is repeated for the remaining service functions until all of the service functions are assigned to resources in the cloud infrastructure.Type: GrantFiled: January 22, 2020Date of Patent: February 24, 2026Assignee: Telefonaktiebolaget LM Ericsson (publ)Inventors: Róbert Szabó, Michael Martin, Malgorzata Svensson, Massimiliano Maggiari, Ákos Recse, Balázs Németh
-
Patent number: 12554527Abstract: Disclosed herein are systems and method for live migration of a guest OS from a source computing device to a target computing device, the method comprising: interrupting execution of the guest OS in a hypervisor on the source computing device, transferring a state of the guest OS to a hypervisor on the target computing device, creating fake I/O requests corresponding to pending I/O requests, resuming execution of the guest OS in the hypervisor on the target computing device without waiting for completion of the pending I/O requests on the source computing device, for each of the pending I/O requests, sending, by the source computing device, a notification about completion of the pending I/O request, and for each pending I/O request for which the notification about completion is received, completing, by the hypervisor on the target computing device, a fake I/O request corresponding to the pending I/O request.Type: GrantFiled: August 30, 2023Date of Patent: February 17, 2026Assignee: Virtuozzo International GmbHInventor: Denis Lunev
-
Patent number: 12517766Abstract: This disclosure relates to a control method and apparatus of cluster resources, and a cloud computing system, and relates to the field of computer technologies. The method includes: in the case where a to-be-controlled resource is a to-be-expanded resource, determining a binding relationship between the to-be-expanded resource and an application; adding the to-be-expanded resource that is initialized into a resource pool of a corresponding application having the binding relationship with the to-be-expanded resource; generating a to-be-executed data packet of a to-be-processed application according to a deployment type of the to-be-processed application; and deploying the to-be-executed data packet on the to-be-expanded resource in the resource pool of the to-be-processed application for execution.Type: GrantFiled: September 24, 2020Date of Patent: January 6, 2026Assignees: BEIJING JINGDONG SHANGKE INFORMATION TECHNOLOGY CO., LTD., BEIJING JINGDONG CENTURY TRADING CO., LTD.Inventors: Bowei Shen, Haifeng Du, Wenqiao Li, Jun Wang, Shi Bai, Chuyi Han
-
Patent number: 12511151Abstract: Disclosure is made of methods, apparatus and system for migrating virtual machines (VMs) between source and destination in a computing environment and, more specifically, to replication based migration. VMs migration is controlled so as to manage transferal of data associated with one or more VMs from a source location to a destination location to meet certain user definable or system constraints. Dynamic control and adjustment of system parameters associated with the migration is also disclosed.Type: GrantFiled: December 12, 2023Date of Patent: December 30, 2025Assignee: Google LLCInventors: Or Igelka, Leonid Vasetsky