Patents by Inventor Irtsam Ghazi

Irtsam Ghazi has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Patent number: 12633297
    Abstract: Methods and systems of processing audio data with a multi-stage audio front end model is provided. A one-dimensional audio waveform is received as input and processed using a multi-stage audio frontend model to convert the one-dimensional waveform into a two-dimensional matrix representing features of the audio waveform. The multi-stage learnable audio frontend model is configured to apply a first filterbank to the audio waveform to generate a first time-frequency representation of the audio waveform; apply a first decimation filter to the audio waveform to generate a first decimated audio input; apply a second filterbank to the first decimated audio input to generate a second time-frequency representation of the audio waveform; and stack the first time-frequency representation and the second time-frequency representation together to generate the two-dimensional matrix.
    Type: Grant
    Filed: September 14, 2023
    Date of Patent: May 19, 2026
    Assignee: Robert Bosch GmbH
    Inventors: Luca Bondi, Irtsam Ghazi, Charles Shelton, Samarjit Das
  • Patent number: 12608422
    Abstract: A method includes defining, by a text encoder, a set of text embeddings for a text prompt indicative of a search query for video content having audio data indicative of a sound feature that is defined as a search parameter of the search query, and ranking a plurality of audio embeddings indicative of a plurality of audio signals and provided in a vector database using the set of text embeddings of the search query. The method further includes detecting a relevant audio record associated with an identified audio embedding from among the ranked audio embeddings, and outputting a relevant video content associated with the relevant audio record to have a computing device play the video content, the relevant video content being obtained from among a plurality of video content.
    Type: Grant
    Filed: July 24, 2024
    Date of Patent: April 21, 2026
    Assignee: Robert Bosch GmbH
    Inventors: Irtsam Ghazi, Ho-Hsiang Wu, Ajit Belsarkar, Luca Bondi, Wei-Cheng Lin, Samarjit Das
  • Patent number: 12586377
    Abstract: Methods and systems for predicting aggressive behavior associated with a surveillance scene. Images are generated from a camera of a surveillance scene, and audio is also generated. A local computing system executes an object classification model on the images to predict one or more classes of objects in the scene. The local computing system also executes a sound-event detection model on the audio to predict one or more classes of events occurring in the scene. Metadata is generated associated with the image-based classes and audio-based classes. The metadata is transferred to a remote computing system that executes a knowledge graph on the metadata to implement knowledge graph-based reasoning to predict aggressive behavior occurring in the surveillance scene based on the metadata. The metadata associated with the predicted aggressive behavior is labeled as such, and control commands are output accordingly.
    Type: Grant
    Filed: December 21, 2023
    Date of Patent: March 24, 2026
    Assignee: Robert Bosch GmbH
    Inventors: Irtsam Ghazi, Alessandro Oltramari
  • Publication number: 20260080895
    Abstract: A method for real-time sound event detection on an embedded device includes pretraining a contrastive language-audio pretraining model as an audio foundation model and preparing offline multimodal query prototypes for sound events of interest. The pretrained model and query prototypes are deployed on an embedded device. The device receives an input audio stream and extracts audio embeddings using the pretrained model. Similarity scores are calculated between the extracted audio embeddings and the prepared query prototypes. The presence of a sound event is determined based on the calculated similarity scores, and a real-time sound event detection result is output. The system includes a memory storing the pretrained model and query prototypes, an audio input interface, and a processor configured to perform the extraction, calculation, determination, and output operations. A non-transitory computer-readable medium stores instructions that, when executed, cause a processor to perform the method.
    Type: Application
    Filed: September 16, 2024
    Publication date: March 19, 2026
    Inventors: Wei-Cheng LIN, Ho-Hsiang WU, Irtsam GHAZI, Luca BONDI, Ajit BELSARKAR, Samarjit DAS
  • Publication number: 20260030294
    Abstract: A method includes defining, by a text encoder, a set of text embeddings for a text prompt indicative of a search query for video content having audio data indicative of a sound feature that is defined as a search parameter of the search query, and ranking a plurality of audio embeddings indicative of a plurality of audio signals and provided in a vector database using the set of text embeddings of the search query. The method further includes detecting a relevant audio record associated with an identified audio embedding from among the ranked audio embeddings, and outputting a relevant video content associated with the relevant audio record to have a computing device play the video content, the relevant video content being obtained from among a plurality of video content.
    Type: Application
    Filed: July 24, 2024
    Publication date: January 29, 2026
    Inventors: Irtsam Ghazi, Ho-Hsiang Wu, Ajit Belsarkar, Luca Bondi, Wei-Cheng Lin, Samarjit Das
  • Patent number: 12449799
    Abstract: Systems and methods for anomaly detection in embedded systems. One example provides an anomaly detection system including a device configured to be connected to a server. The device includes an electronic processor and a memory. The memory stores a first application. The electronic processor is configured to transmit a request for an anomaly detection application. The anomaly detection application includes a set of instructions for detecting anomalies of the device. The electronic processor is configured to receive, from the server, the anomaly detection application, and replace the first application with the anomaly detection application within the memory.
    Type: Grant
    Filed: November 7, 2023
    Date of Patent: October 21, 2025
    Assignee: Robert Bosch GmbH
    Inventors: Irtsam Ghazi, Emily Ruppel, Shruti Lall
  • Patent number: 12437774
    Abstract: Systems and methods for audio event detection. One example provides an event detection system comprising a plurality of audio devices. Each of the plurality of audio devices is configured to be communicatively coupled to a server and includes an electronic processor. The electronic processor is configured to detect, via a microphone, audio, and determine an audio event within the audio. The electronic processor is configured to receive an image from a camera, associate the image data with the audio event to generate event metadata, and transmit the event metadata, the audio, and the image to the server.
    Type: Grant
    Filed: November 9, 2022
    Date of Patent: October 7, 2025
    Assignee: Robert Bosch GmbH
    Inventors: Ajit Belsarkar, Jacob A. Gallucci, Irtsam Ghazi
  • Publication number: 20250209819
    Abstract: Methods and systems for predicting aggressive behavior associated with a surveillance scene. Images are generated from a camera of a surveillance scene, and audio is also generated. A local computing system executes an object classification model on the images to predict one or more classes of objects in the scene. The local computing system also executes a sound-event detection model on the audio to predict one or more classes of events occurring in the scene. Metadata is generated associated with the image-based classes and audio-based classes. The metadata is transferred to a remote computing system that executes a knowledge graph on the metadata to implement knowledge graph-based reasoning to predict aggressive behavior occurring in the surveillance scene based on the metadata. The metadata associated with the predicted aggressive behavior is labeled as such, and control commands are output accordingly.
    Type: Application
    Filed: December 21, 2023
    Publication date: June 26, 2025
    Inventors: Irtsam Ghazi, Alessandro Oltramari
  • Patent number: 12293773
    Abstract: A system for automatically selecting a sound recognition model for an environment based on audio data and image data associated with the environment. The system includes a camera, a microphone, a memory including a plurality of sound recognition models, and an electronic processor. The electronic processor is configured to receive the audio data associated with the environment from the microphone, receive the image data associated with the environment from the camera, and determine one or more characteristics of the environment based on the audio data and the image data. The electronic processor is also configured to select the sound recognition model from the plurality of sound recognition models based on the one or more characteristics of the environment, receive additional audio data associated with the environment from the microphone, and analyze the additional audio data using the sound recognition model to perform a sound recognition task.
    Type: Grant
    Filed: November 3, 2022
    Date of Patent: May 6, 2025
    Assignee: Robert Bosch GmbH
    Inventors: Luca Bondi, Irtsam Ghazi
  • Publication number: 20250095664
    Abstract: Methods and systems of processing audio data with a multi-stage audio front end model is provided. A one-dimensional audio waveform is received as input and processed using a multi-stage audio frontend model to convert the one-dimensional waveform into a two-dimensional matrix representing features of the audio waveform. The multi-stage learnable audio frontend model is configured to apply a first filterbank to the audio waveform to generate a first time-frequency representation of the audio waveform; apply a first decimation filter to the audio waveform to generate a first decimated audio input; apply a second filterbank to the first decimated audio input to generate a second time-frequency representation of the audio waveform; and stack the first time-frequency representation and the second time-frequency representation together to generate the two-dimensional matrix.
    Type: Application
    Filed: September 14, 2023
    Publication date: March 20, 2025
    Inventors: Luca BONDI, Irtsam GHAZI, Charles SHELTON, Samarjit DAS
  • Publication number: 20250076867
    Abstract: Systems and methods for anomaly detection in embedded systems. One example provides an anomaly detection system including a device configured to be connected to a server. The device includes an electronic processor and a memory. The memory stores a first application. The electronic processor is configured to transmit a request for an anomaly detection application. The anomaly detection application includes a set of instructions for detecting anomalies of the device. The electronic processor is configured to receive, from the server, the anomaly detection application, and replace the first application with the anomaly detection application within the memory.
    Type: Application
    Filed: November 7, 2023
    Publication date: March 6, 2025
    Inventors: Irtsam Ghazi, Emily Ruppel, Shruti Lall
  • Publication number: 20240153526
    Abstract: Systems and methods for audio event detection. One example provides an event detection system comprising a plurality of audio devices. Each of the plurality of audio devices is configured to be communicatively coupled to a server and includes an electronic processor. The electronic processor is configured to detect, via a microphone, audio, and determine an audio event within the audio. The electronic processor is configured to receive an image from a camera, associate the image data with the audio event to generate event metadata, and transmit the event metadata, the audio, and the image to the server.
    Type: Application
    Filed: November 9, 2022
    Publication date: May 9, 2024
    Inventors: Ajit Belsarkar, Jacob A. Gallucci, Irtsam Ghazi
  • Publication number: 20240153524
    Abstract: A system for automatically selecting a sound recognition model for an environment based on audio data and image data associated with the environment. The system includes a camera, a microphone, a memory including a plurality of sound recognition models, and an electronic processor. The electronic processor is configured to receive the audio data associated with the environment from the microphone, receive the image data associated with the environment from the camera, and determine one or more characteristics of the environment based on the audio data and the image data. The electronic processor is also configured to select the sound recognition model from the plurality of sound recognition models based on the one or more characteristics of the environment, receive additional audio data associated with the environment from the microphone, and analyze the additional audio data using the sound recognition model to perform a sound recognition task.
    Type: Application
    Filed: November 3, 2022
    Publication date: May 9, 2024
    Inventors: Luca Bondi, Irtsam Ghazi