Patents Examined by Gabriel Chu
-
Patent number: 12675354Abstract: A system is described having one or more processing devices that prepare a work queue entry (WQE), transmit the WQE or provide a doorbell indication indicating the WQE is available, and, in response to determining an execution failure has occurred, perform at least one additional step to facilitate completion of an operation associated with the WQE.Type: GrantFiled: October 10, 2024Date of Patent: July 7, 2026Assignee: NVIDIA CorporationInventors: Pak Markthub, Daniel Marcovitch, Akhil Langer, Yossef Itigin
-
Patent number: 12639179Abstract: Techniques are provided for metadata management for enabling automated switchover in accordance with a configuration of storage solution that expresses a preference for either maintaining availability (e.g., a non-zero RPO mode) of the storage solution or avoiding data loss (e.g., a zero RPO mode). In one example, responsive to detecting a switchover trigger event, a node of a local cluster of a cross-site storage solution determines whether performance of an automated switchover from a failed cluster to a surviving cluster of the cross-site storage solution is enabled. Responsive to an affirmative determination, the node selectively proceeds with the automated switchover based on the configuration.Type: GrantFiled: July 8, 2024Date of Patent: May 26, 2026Assignee: NETAPP, INC.Inventors: Sasidharan Krishnan, Kalaivani Arumugham, Preksha Bansal, Vijay Kumar Chakravarthy Ekkaladevi, Ryan Edward Bartlett
-
Patent number: 12625782Abstract: A Network Virtualization Device (NVD) executes a set of Virtual Network Interface Cards (VNICs). The set of VNICs includes a first VNIC that forwards packets for a set of one or more packet flows. The NVD stores a first VNIC-related information that includes information identifying a first set of one or more packet flows and associated state information The NVD in response to determining that the state information for the first VNIC is to be synchronized with another NVD, identifies a first backup NVD for the first VNIC, wherein the first backup NVD is a backup for the first VNIC, and communicates to the first backup NVD, a portion of the state information stored by the NVD for the first VNIC.Type: GrantFiled: November 18, 2024Date of Patent: May 12, 2026Assignee: ORACLE INTERNATIONAL CORPORATIONInventors: Jagwinder Singh Brar, Eugene Nalimov, Steven Chervets, Abhay Patil, Michal Aleksander Karczmarek
-
Patent number: 12608268Abstract: Methods and systems for managing operation of a deployment comprising data processing systems are disclosed. The operation may be managed by optimizing a performance of a data processing system. The optimization may include remediation of the data processing system due to an occurrence of an issue, disruption, etc. To optimize the performance, the data processing system may perform a similarity analysis. During the similarity analysis, the data processing system may use a similarity map to search for at least one other data processing system that is similar to the data processing system. The data processing system may collaborate with the at least one other data processing system to identify a remediation procedure for the issue, the disruption, etc. The data processing system may then perform the remediation procedure.Type: GrantFiled: December 20, 2024Date of Patent: April 21, 2026Assignee: Dell Products L.P.Inventors: Ashok Narayanan Potti, Zijia Wang, Tsehsin Jason Liu, Dale Wang, Min Gong
-
Patent number: 12602289Abstract: A method may include: a listener application receiving a health status of cloud applications in cloud pools, each cloud pool hosted by cloud hardware; the listener application receiving a maintenance schedule comprising a plurality of maintenance events for the cloud application and/or the cloud hardware, the maintenance schedule; the listener application identifying an impacted cloud application and/or cloud hardware that is scheduled for one of the maintenance events; the listener application mapping the impacted cloud hardware and/or cloud applications to a global load balancer; a remediation application sending a first command to the mapped global load balancer to disable traffic to the impacted cloud hardware and/or cloud applications; the remediation application executing the maintenance events on the cloud hardware and/or cloud applications; the remediation application sending a second command to the global mapped global load balancer to re-enable traffic to the impacted cloud hardware and/or cloud appType: GrantFiled: November 18, 2024Date of Patent: April 14, 2026Assignee: JPMORGAN CHASE BANK, N.A.Inventors: Ashish Lamichhane, Kiran Sekharamahanthi, Allie Pistolessi, Robert Huarcaya Lazaro
-
Patent number: 12591481Abstract: A system and method for anomaly prediction in a plurality of applications to maintain operability of the system. Real-time data from the plurality of applications can be received, and for each desired variable in the data an entropy that can indicate an anomaly in the respective data. The data can be divided into batches of fixes size and a plurality of machine learning (ML) models applied to each of the batches to determine whether the data contains an anomaly or is about to contain an anomaly. The plurality of ML models can be trained based on historical data from the plurality of applications. A large language model (LLM) can be applied to the output of the plurality of ML models to evaluate performance/fine tune the plurality of ML models, and an alert can be transmitted to a system that automatically executes script(s) to take action to avoid the anomaly.Type: GrantFiled: July 15, 2025Date of Patent: March 31, 2026Assignee: Morgan Stanley Services Group Inc.Inventors: Vincent Petrillo, Afreen Bano Ansari, Soham Sen, Tarik M. Sheth, Harsh Sharma
-
Patent number: 12585517Abstract: One example method includes monitoring confidence scores received from an edge entity in an edge environment, and the edge entity comprises hardware and/or software, comparing a change in the confidence scores with a threshold permissible confidence score change, when a magnitude of the change is above a threshold permissible change, identifying the edge entity, or a system/device monitored by the edge entity, as a possibly malfunctioning entity, identifying a corrective action for the edge entity, and implementing, and/or causing implementation of, the corrective action.Type: GrantFiled: June 21, 2024Date of Patent: March 24, 2026Assignee: Dell Products L.P.Inventors: Pankaj Pande, Stephen J. Todd
-
Patent number: 12572402Abstract: Systems, methods, and techniques described herein provide timeout error detection in computing systems. In an aspect, a requesting entity transmits, in a first time interval and on behalf of a first queue, a first request to a responding entity. Responsive to transmitting the first request, the requesting entity increments a value of a count. The requesting entity transmits, in the first interval and on behalf of a second queue, a second request to a responding entity. Responsive to transmitting the second request, the requesting entity increments the value of the count. Responsive to receiving a response from the responding entity prior to expiration of a second time interval, the requesting entity decrements the value of the count. Subsequent to expiration of the second time interval, the requesting entity determines whether or not the value of the count is nonzero. If so, the requesting entity causes performance of a timeout action.Type: GrantFiled: October 4, 2024Date of Patent: March 10, 2026Assignee: MICROSOFT TECHNOLOGY LICENSING, LLCInventors: Dongwook Lee, Galina Vladimirovna Malakhova, Tama Gal, Yevgeny Yankilevich, Mahmoud El Haddad
-
Patent number: 12566682Abstract: Systems and methods are provided for failure resiliency in distributed training of machine learning (ML) models. Examples include a plurality of compute nodes storing shards of a plurality of shards of model states of an ML model, and a first compute node storing a first shard of model states of the ML model. The first compute node can store a plurality of shard portions. Each shard portion can be received from a respective compute node of the plurality of compute nodes and can be a replica of a portion of a respective shard, of the plurality of shards, stored at the respective compute node. Responsive to a failure of a compute node of the plurality of compute nodes, the first compute node can update the first shard with a shard portion corresponding to the failed compute node and the ML model can be trained based on the updated first shard.Type: GrantFiled: April 12, 2024Date of Patent: March 3, 2026Assignee: Hewlett Packard Enterprise Development LPInventors: Lianjie Cao, Saeed Rashidi, Puneet Sharma, Garrett Goon, Paolo Faraboschi
-
Patent number: 12524289Abstract: A method for execution by a computing device of a storage network includes obtaining performance information for a storage device of a set of storage devices of the storage network, wherein data is error encoded into sets of encoded data slices that are stored in the storage devices. The method further includes obtaining additional performance information for each storage device of the storage devices, wherein the additional performance information is based on historical data. The method further includes comparing the performance information to the additional performance information to produce comparison performance information and identifying at least one component of the comparison performance information. The method further includes comparing the at least one component to a corresponding error threshold and outputting indication of a performance error for the storage device when at least one component of the comparison performance information is greater than an error threshold.Type: GrantFiled: February 21, 2023Date of Patent: January 13, 2026Assignee: Pure Storage, Inc.Inventors: Greg R. Dhuse, Jason K. Resch, Ilya Volvovski
-
Patent number: 12524319Abstract: Systems and methods are provided for failure resiliency in distributed training of machine learning (ML) models. Examples include a plurality of compute nodes storing optimizer shards of a plurality of optimizer shards and a first compute node storing a first optimizer shard of optimizer states. The first compute node can store optimizer shard portions, each of which can be received from a respective compute node of the plurality of compute nodes and can be a replica of a portion of a respective optimizer shard of the plurality of optimizer shards, stored at the respective compute node. Responsive to a failure of a compute node of the plurality of compute nodes, the first compute node can update the first optimizer shard with an optimizer shard portion corresponding to the failed compute node and the ML model can be trained based on the updated first optimizer shard.Type: GrantFiled: April 26, 2024Date of Patent: January 13, 2026Assignee: Hewlett Packard Enterprise Development LPInventors: Lianjie Cao, Saeed Rashidi, Garrett Goon, Paolo Faraboschi, Puneet Sharma
-
Patent number: 12515681Abstract: A method of runtime program verification receives, through a preemptive mechanism of an operating system running on an external system, recent values for a set of system health indicators for the external system. The recent values are compared to calibrated values for the set of system health indicators, and normal operation of the external system is verified based on whether the comparison of the recent values to the calibrated values are within predetermined error thresholds.Type: GrantFiled: June 9, 2023Date of Patent: January 6, 2026Assignee: Mercedes-Benz Group AGInventor: Francois Piednoel
-
Patent number: 12505020Abstract: An arithmetic operation target is input to an information processing apparatus which causes an accelerator apparatus to perform an arithmetic operation using the arithmetic operation target. The information processing apparatus performs, regarding each of a plurality of arithmetic operation elements of the arithmetic operation target, whether to allocate one or more diagnostic circuits, which are available processing circuits for an accuracy diagnosis of the arithmetic operation from a plurality of processing circuits in the accelerator apparatus to the arithmetic operation element on the basis of a failure influence degree.Type: GrantFiled: September 26, 2022Date of Patent: December 23, 2025Assignee: HITACHI, LTD.Inventors: Takumi Uezono, Hiroaki Itsuji, Kenichi Shimbo, Masayoshi Takahashi, Yutaka Uematsu
-
Patent number: 12493511Abstract: A method is disclosed for alerting a co-tenant of a fault in a partitioned system. An apparatus and system also perform the functions of the method. The method includes monitoring, for faults, two or more partitions in a system sharing a management controller where each of the partitions is associated with a tenant. The method includes detecting a fault on a first partition and identifying an alert class of multiple alert classes for the fault. The method includes identifying the alert class of the fault as a high-level alert class indicating that the fault affects one or more other partitions different than the first partition. The method includes notifying the tenant of the first partition and each tenant of the one or more other partitions of the fault in response to the alert class being the high-level alert class.Type: GrantFiled: March 5, 2024Date of Patent: December 9, 2025Assignee: Lenovo Enterprise Solutions (Singapore) Pte. Ltd.Inventors: Gary D. Cudak, Pravin S. Patel, James Parsonese, Michael J. Biersack, Mehul Shah
-
Patent number: 12481568Abstract: A data storage infrastructure may establish a partition that includes a first data center and a second data center that is geographically separated from the first data center. The data storage infrastructure may replicate a full snapshot and one or more incremental snapshots of a virtual machine from a first data management platform to a second data management platform, where the virtual machine is migrated from a first host of the first host group to a second host of the second host group upon a failover event occurring at the first data center. The data storage infrastructure may then capture an incremental snapshot of the virtual machine based on linking a first instance of the virtual machine that was replicated from the first data management platform and a second instance of the virtual machine that is managed by the second data management platform.Type: GrantFiled: January 25, 2024Date of Patent: November 25, 2025Assignee: Rubrik, Inc.Inventors: Disheng Su, Bharadwaj Rayala, Li Ding
-
Patent number: 12461805Abstract: Described are various embodiments of system and method for monitoring a website. The system and method allows a website administrator to monitor for specific user event sequences or actions that will trigger a page check to see if one or more web elements or features are correctly displayed or not. In one embodiment, the method comprises the step of monitoring, by a first processor of each of one or more user devices visiting the website, one or more user events. From the one or more user events, the presence of one or more pre-conditions is confirmed, upon the one or more pre-conditions being identified, the method assess the validity of one or more assertions. A failure rate of a number of assertions is checked and compared to a designated threshold. If the designated threshold is exceeded, a notification is sent.Type: GrantFiled: May 2, 2024Date of Patent: November 4, 2025Inventors: Joshua Koopferstock, Reid Tait, David Seel, Kiril Kuts, Oleksii Babik, Filip Slatinac, Patrick Soutar, Viktor Musiienko
-
Patent number: 12443478Abstract: A computer system performs tasks in an access restricted environment. Data is logged in diagnostic files about logical resources in use by the computer system as the computer system attempts to perform the tasks. Occasionally, a problem may prevent the computer system from correctly performing a task. A machine authenticates a user in the access-restricted environment and receives error metadata to initiate an automated process for generating a troubleshooting signature. The automated process involves selecting a metadata extraction policy based on a category of error and using the metadata extraction policy to extract metadata from a diagnostic file. The extracted metadata is analyzed to determine troubleshooting components including the problem, a source of the problem, and/or a version of software that encountered the problem.Type: GrantFiled: January 16, 2024Date of Patent: October 14, 2025Assignee: Oracle International CorporationInventors: Nagarajan Muthukrishnan, Srikanth Nagandla, Yu Li, Paul Hsu
-
Patent number: 12430189Abstract: According to some embodiments, systems and methods are provided including a memory storing processor-executable program code; and a processing unit to execute the processor-executable program code to cause the system to: receive a service disruption notification for a service; identify a service disruption type based on the received service disruption notification; generate disruption identification instructions in response to the identified service disruption type; display the generated disruption identification instructions on a user interface; receive an action command in a user entry field of the user interface, the action command including a service name of the service; and dynamically generate a response to the received action command. Numerous other aspects are provided.Type: GrantFiled: November 16, 2023Date of Patent: September 30, 2025Assignee: SAP SEInventors: Aleksandar Gospodinov Gagov, Svetoslav Dimitrov Milkov, Nikolay Feodorov Kitanov, Georgi Veselinov Georgiev
-
Patent number: 12430223Abstract: An embodiment includes detecting an application availability metric and a performance data metric of a system by a Tracing Policy Analyzer, responsive to detecting the application availability metric and the performance data metric, training a Bayesian optimization model of the Tracing Policy Analyzer based on the application availability metric and the performance data metric where the Bayesian optimization model outputs an optimal tracing policy. The embodiment includes detecting the optimal tracing policy by a Tracing Level Determination Module of an application of the system, responsive to detecting the optimal tracing policy, computing a computed tracing level by the Tracing Level Determination Module executing a k-nearest neighbors algorithm based on the optimal tracing policy. The embodiment also includes tracing of the application of the system by a Tracer based on the computed tracing level.Type: GrantFiled: January 8, 2024Date of Patent: September 30, 2025Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATIONInventors: Yong Hu Sun, Xiao Juan Niu, Li Jian Wang, Pei Ran Han, Xing Tian, Shao Rong Li
-
Patent number: 12430198Abstract: In an illustrative embodiment, systems and methods for predicting network service outages gather log data from processes executing on computing device(s), produce training data from the log data for training tree-based machine learning model(s) and metrics calculated therefrom, periodically apply data sets derived from future log data and metrics calculated therefrom to the machine learning model(s) to predict critical error(s)/system outage(s), and periodically update the trained machine learning models using the periodically applied data sets.Type: GrantFiled: October 6, 2022Date of Patent: September 30, 2025Assignee: Federal Home Loan Mortgage Corporation (“FREDDIE MAC”)Inventors: Keenan Moukarzel, Tatyana Krol, Chengcheng Xiong