Abstract: The present invention provides a system and method for real-time target speaker recognition and transcription, enhancing accuracy in noisy and multi-speaker environments. Using audio and video inputs from a microphone and camera, the system processes features through neural networks, including cross-attention mechanisms, to isolate speaker-specific data and suppress noise. A Speaker Module tracks and identifies active speakers via face detection, recurrent neural networks, or decision mechanisms, ensuring transcription continuity even with temporary visual occlusions. Outputs generated by the invention include transcribed text and speaker labels, displayed on an augmented reality display or a mobile device as real-time subtitles. The system can suppress the display of the user's own speech for improved usability and supports multilingual transcription. This invention offers a robust solution for individuals with hearing impairments, enabling accurate comprehension in challenging acoustic settings.
Type:
Application
Filed:
February 4, 2025
Publication date:
August 6, 2026
Applicant:
RTC Vision Ltd.
Inventors:
Ron WEIN, Joseph RUBNER, Nir AVRAHAMI, Nir TSUK, Kfir GEDALYAHU
Abstract: Systems and methods are provided for recognizing characters within a distorted image. According to a one aspect, a method for recognizing one or more characters within a distorted image includes rendering one or more imitation images, the imitation images including simulations of the distorted image, applying one or more distortion models to the imitation images, thereby generating distorted imitation images, comparing the distorted imitation images with the distorted image in order to compute similarities between the distorted imitation images and the distorted image, and identifying the characters based on the best similarity. According to other aspects, the systems and methods can be configured to provide recognition of other distorted data types and elements.
Type:
Grant
Filed:
January 18, 2012
Date of Patent:
March 24, 2015
Assignee:
RTC Vision Ltd.
Inventors:
Sefy Kagarlitsky, Joseph Rubner, Nir Avrahami, Yohai Falik
Abstract: Systems and methods are disclosed for enhancing digital signals. In one implementation, a digital signal that has undergone a non-linear distortion can be received. The non-linear distortion can be reformulated as one or more linear operators that yield a statistical connection between a first signal and a second signal and one or more convex constraints on the first signal and/or the second signal. A convex minimization problem can be formulated in view of the first signal, the second signal, and the one or more convex constraints. The digital signal can be processed to solve the convex minimization problem, thereby generating an enhanced digital signal.
Abstract: Systems and methods are provided for recognizing characters within a distorted image. According to a one aspect, a method for recognizing one or more characters within a distorted image includes rendering one or more imitation images, the imitation images including simulations of the distorted image, applying one or more distortion models to the imitation images, thereby generating distorted imitation images, comparing the distorted imitation images with the distorted image in order to compute similarities between the distorted imitation images and the distorted image, and identifying the characters based on the best similarity. According to other aspects, the systems and methods can be configured to provide recognition of other distorted data types and elements.
Type:
Application
Filed:
January 18, 2012
Publication date:
October 24, 2013
Applicant:
RTC VISION LTD.
Inventors:
Sefy Kagarlitsky, Joseph Rubner, Nir Avrahami, Yohai Falik