Abstract: In various embodiments, a process for teaching computer-use artificial intelligence agents includes receiving a task recording including video information that captures a multi-step interaction with a user interface and associated device input information. A sequence of anchor actions is extracted from the video information. For each anchor action, a preceding frame and a following frame are obtained adjacent to the anchor action. A prompt including a task prompt in natural language that describes the interaction based on the anchor actions and frames is generated. The task prompt includes computer-executable instructions to replicate the interaction. The prompt is executed to obtain an execution trace including predicted actions. The predicted actions are compared with the anchor actions to identify whether there is a discrepancy between the predicted and anchor actions. If there is a discrepancy, the prompt is iteratively enhanced until there is no discrepancy.
Type:
Grant
Filed:
December 23, 2025
Date of Patent:
August 25, 2026
Assignee:
OPNOVA, Inc.
Inventors:
Pedro dos Santos Saleiro, Sinan Eren, Vladimir Balayan, Jose Luis Ferras Pereira, Tiago Altino de Andrade e Melo