Abstract: A procedure generation system obtains multimedia content describing performance of a task and generates a procedure including content for guiding a user through performance of the task. The procedure generation system extracts audio data from the multimedia content and generates a transcription of the audio data through application of a trained model. The transcription includes text corresponding to the audio data and timestamps associated with different text. Based on the transcription, a trained model generates a set of steps, with each step including text corresponding to different time intervals. The procedure generation system identifies portions of the multimedia content corresponding to different steps based on the time intervals and associates identified portions of the multimedia content with corresponding steps to generate the procedure. This generates a procedure with various steps including text and a corresponding portion of the multimedia content.
Type:
Grant
Filed:
May 23, 2024
Date of Patent:
August 18, 2026
Assignee:
Squint, Inc.
Inventors:
Devin Bhushan, Dylan Patricia Conway, Benjamin Scott Weaver, Jim Zhu
Abstract: Aspects of the present disclosure are directed to using AI tools such as large language models, grounded computer vision models, and visual-language models to verify that the state of a target is the desired state of the target. An image of the target and a semantic textual description of the desired state of the target may be used to determine whether the state of the target is the desired state.
Type:
Grant
Filed:
August 28, 2025
Date of Patent:
March 3, 2026
Assignee:
Squint, Inc.
Inventors:
Jacob Hicks, Grant Gliner, Devin Bhushan, Alexander Young, Jim Zhu, Benjamin Weaver, Dane Laughlin