Patents by Inventor Muhammad MAAZ

Muhammad MAAZ has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20260017926
    Abstract: A system and method for grounded multimodal conversation in which a global image encoder is connected to a vision-to-language (V-L) projection layer for encoding an image and projecting the image into scene text. A region encoder constructs a feature pyramid from layers of the global image encoder, followed by a Region of Interest layer to generate a feature map. The V-L projection layer maps features into projected image features. A large language model receives an input of an augmentation of text instruction and features and generates a conversation concerning the image. A language-to-prompt projection layer transforms last-layer embeddings of the large language model corresponding to segment tokens into a pixel space then a pixel decoder utilizes the pixel feature space together with a grounding image encoder to produce pixel-level object grounding.
    Type: Application
    Filed: July 15, 2024
    Publication date: January 15, 2026
    Applicant: Mohamed bin Zayed University of Artificial Intelligence
    Inventors: Hanoona BANGALATH, Muhammad MAAZ, Sahal Shaji MULLAPPILLY, Abdelrahman YOUSSIEF, Salman KHAN, Hisham CHOLAKKAL, Rao ANSWER, Eric XING, Fahad S. KHAN
  • Patent number: 12493741
    Abstract: A method and system for multi-modal prompt learning of vision-language models. Encodings of image-text pairs can be combined with image prompts and text prompts before being input into an image encoder and text encoder of a vision-language model respectively. The image prompt can be generated using the text prompt using a vision-language coupling function to encourage synergy between the two prompts. The combination of encodings and prompts can be fed through the transformer layers of the encoders, and the output of each layer can be combined with a new prompt before entering the next layer, up until a specific depth. The subsequent transformer layers can process the output and generate a final representation for the image and text which can then be used for downstream tasks.
    Type: Grant
    Filed: December 28, 2022
    Date of Patent: December 9, 2025
    Assignee: Mohamed bin Zayed University of Artificial Intelligence
    Inventors: Muhammad Uzair Khattak, Hanoona Abdul Rasheed Bangalath, Muhammad Maaz, Salman Khan, Fahad Shahbaz Khan
  • Patent number: 12373956
    Abstract: A system for 3D medical image segmentation includes a medical imaging device for obtaining a plurality of 2D images forming a volumetric image, processing circuitry, and a display. The processing circuitry is configured with a first stage to divide the volumetric image into 3D image patches, a hierarchical encoder-decoder structure in which resolution of features of the 3D image patches is decreased by a factor of two in each of a plurality of stages of the encoder, an encoder output connected to the decoder via skip connections, and a convolutional block to produce a voxel-wise final segmentation mask. The encoder includes a plurality of efficient paired attention blocks each with a spatial attention branch and a channel attention branch that learn respective spatial and channel attention feature maps. The display displays the final segmentation mask.
    Type: Grant
    Filed: April 26, 2023
    Date of Patent: July 29, 2025
    Assignee: Mohamed bin Zayed University of Artificial Intelligence
    Inventors: Abdelrahman Shaker, Muhammad Maaz, Hanoona Rasheed, Salman Khan, Fahad Shahbaz Khan
  • Patent number: 12373672
    Abstract: An edge computing system, computer readable storage medium and method for object detection, including processing circuitry. The processing circuitry is configured with a hybrid CNN and vision transformer backbone network in an object detection deep learning network. The backbone network receives an image, and includes a first convolutional encoder to extract local features from feature maps of the image, a second stage having consecutive second convolutional encoders, a positional encoding layer, split depth-wise transpose attention (SDTA) encoders, consecutive convolutional encoders, a third stage and a fourth stage SDTA encoder. Each of the SDTA encoders perform multi-headed self-attention by applying a dot product operation across channel dimensions in order to compute cross-covariance across channels to generate attention feature maps.
    Type: Grant
    Filed: December 9, 2022
    Date of Patent: July 29, 2025
    Assignee: Mohamed bin Zayed University of Artificial Intelligence
    Inventors: Muhammad Maaz, Abdelrahman Shaker, Hisham Cholakkal, Salman Khan, Syed Waqas Zamir, Rao Muhammad Anwer, Fahad Shahbaz Khan
  • Patent number: 12288372
    Abstract: An object detection system and method in which a machine learning engine is configured with a region-based knowledge distillation stage that generates region embeddings from a training image having bounding boxes. A linear layer learns a region-level vision-language mapping for projecting feature embeddings from the training image to a common feature space shared by text embeddings to obtain the region embeddings. An image-level supervision stage generates pseudo-box labels for a classification training image and region embeddings from the training image having bounding boxes and corresponding class labels and the classification training image having an image-level label as input. Pseudo-box labels are determined on the classification training image as an image-level vision-language mapping. A weight transfer function conditions the image-level vision-language mapping on the learned region-level vision-language mapping.
    Type: Grant
    Filed: December 20, 2022
    Date of Patent: April 29, 2025
    Assignee: Mohamed bin Zayed University of Artificial Intellegence
    Inventors: Hanoona Abdul Rasheed Bangalath, Muhammad Maaz, Muhammad Uzair Khattak, Salman Khan, Fahad Shahbaz Khan
  • Publication number: 20240362788
    Abstract: A system for 3D medical image segmentation includes a medical imaging device for obtaining a plurality of 2D images forming a volumetric image, processing circuitry, and a display. The processing circuitry is configured with a first stage to divide the volumetric image into 3D image patches, a hierarchical encoder-decoder structure in which resolution of features of the 3D image patches is decreased by a factor of two in each of a plurality of stages of the encoder, an encoder output connected to the decoder via skip connections, and a convolutional block to produce a voxel-wise final segmentation mask. The encoder includes a plurality of efficient paired attention blocks each with a spatial attention branch and a channel attention branch that learn respective spatial and channel attention feature maps. The display displays the final segmentation mask.
    Type: Application
    Filed: April 26, 2023
    Publication date: October 31, 2024
    Applicant: Mohamed bin Zayed University of Artificial Intelligence
    Inventors: Abdelrahman SHAKER, Muhammad MAAZ, Hanoona RASHEED, Salman KHAN, Fahad Shahbaz KHAN
  • Publication number: 20240220722
    Abstract: A method and system for multi-modal prompt learning of vision-language models. Encodings of image-text pairs can be combined with image prompts and text prompts before being input into an image encoder and text encoder of a vision-language model respectively. The image prompt can be generated using the text prompt using a vision-language coupling function to encourage synergy between the two prompts. The combination of encodings and prompts can be fed through the transformer layers of the encoders, and the output of each layer can be combined with a new prompt before entering the next layer, up until a specific depth. The subsequent transformer layers can process the output and generate a final representation for the image and text which can then be used for downstream tasks.
    Type: Application
    Filed: December 28, 2022
    Publication date: July 4, 2024
    Applicant: Mohamed bin Zayed University of Artificial Intelligence
    Inventors: Muhammad Uzair KHATTAK, Hanoona Abdul Rasheed BANGALATH, Muhammad MAAZ, Salman KHAN, Fahad Shahbaz KHAN
  • Publication number: 20240203085
    Abstract: An object detection system and method in which a machine learning engine is configured with a region-based knowledge distillation stage that generates region embeddings from a training image having bounding boxes. A linear layer learns a region-level vision-language mapping for projecting feature embeddings from the training image to a common feature space shared by text embeddings to obtain the region embeddings. An image-level supervision stage generates pseudo-box labels for a classification training image and region embeddings from the training image having bounding boxes and corresponding class labels and the classification training image having an image-level label as input. Pseudo-box labels are determined on the classification training image as an image-level vision-language mapping. A weight transfer function conditions the image-level vision-language mapping on the learned region-level vision-language mapping.
    Type: Application
    Filed: December 20, 2022
    Publication date: June 20, 2024
    Applicant: Mohamed bin Zayed University of Artificial Intelligence
    Inventors: Hanoona Abdul Rasheed BANGALATH, Muhammad MAAZ, Muhammad Uzair KHATTAK, Salman KHAN, Fahad Shahbaz KHAN
  • Publication number: 20240193404
    Abstract: An edge computing system, computer readable storage medium and method for object detection, including processing circuitry. The processing circuitry is configured with a hybrid CNN and vision transformer backbone network in an object detection deep learning network. The backbone network receives an image, and includes a first convolutional encoder to extract local features from feature maps of the image, a second stage having consecutive second convolutional encoders, a positional encoding layer, split depth-wise transpose attention (SDTA) encoders, consecutive convolutional encoders, a third stage and a fourth stage SDTA encoder. Each of the SDTA encoders perform multi-headed self-attention by applying a dot product operation across channel dimensions in order to compute cross-covariance across channels to generate attention feature maps.
    Type: Application
    Filed: December 9, 2022
    Publication date: June 13, 2024
    Applicant: Mohamed bin Zayed University of Artificial Intelligence
    Inventors: Muhammad MAAZ, Abdelrahman SHAKER, Hisham CHOLAKKAL, Salman KHAN, Syed Waqas ZAMIR, Rao Muhammad ANWER, Fahad Shahbaz KHAN