Patents by Inventor Ramin Mehran

Ramin Mehran has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20260089368
    Abstract: A method includes segmenting a source video into source video segments, and generating a script for a new video using a generative artificial intelligence (AI) engine. The script includes, for each of one or more new video segments arranged according to a sequential order, a segment descriptor and a segment voice-over transcript. For each new video segment, a voice-over segment is generated from among the source video segments based on the respective segment voice-over transcript, and a set of source video segment(s) is selected based on the respective segment descriptor, for use in generating the new video segment. The method also includes generating the new video, at least in part by inserting the generated voice-over segments for the new video segment(s), and the selected set(s) of source video segment(s) for the new video segment(s), in accordance with the sequential order.
    Type: Application
    Filed: September 25, 2025
    Publication date: March 26, 2026
    Inventors: Zhixian Yu, Bo Hu, Chun-Te Chu, Ramin Mehran, Yukun Zhu, Ying Ding, Shushan Chen, Jiashi Cao, Sudheendra Vijayanarasimhan
  • Publication number: 20260089366
    Abstract: A method for generating a new video from a source video includes determining that the source video is associated with one or more components, and identifying a source starting segment within the source video at least in part by selecting a segment identification model, from among a plurality of candidate segment identification models, based at least in part on the segment identification module being configured to operate upon at least one of the one or more components. The method also includes identifying the source starting segment by using the selected segment identification model to process at least a portion of the source video. The method also includes generating the new video using one or more portions of the source video, wherein generating the new video includes generating an initial segment of the new video based on the source starting segment.
    Type: Application
    Filed: September 30, 2024
    Publication date: March 26, 2026
    Inventors: Zhixian Yu, Bo Hu, Chun-Te Chu, Ramin Mehran, Yukun Zhu, Ying Ding, Shushan Chen, Jiashi Cao, Sudheendra Vijayanarasimhan
  • Publication number: 20260064972
    Abstract: For each portion of a content item, the portion of the content item can be processed with a machine-learned Large Foundational Model (LFM) to obtain an attentional value output comprising a summarization of the portion. The attentional value output can be processed with the machine-learned LFM to obtain an attentional query output descriptive of thematic elements associated with the portion, and an attentional key output comprising key words and/or phrases from the portion. An attentional weight can be determined for each portion based on a semantic similarity between the attentional query output and the attentional key output for each portion. A subset of portions of the content item can be selected based on the attentional weight determined for the subset. A task output can be generated based on the attentional value output obtained for each of the subset of portions.
    Type: Application
    Filed: August 27, 2024
    Publication date: March 5, 2026
    Inventors: Ramin Mehran, Nilpa Jha
  • Patent number: 12154547
    Abstract: A method for multi-channel voice activity detection includes receiving a sequence of input frames characterizing streaming multi-channel audio captured by an array of microphones. Each channel of the streaming multi-channel audio includes respective audio features captured by a separate dedicated microphone. The method also includes determining, using a location fingerprint model, a location fingerprint indicating a location of a source of the multi-channel audio relative to the user device based on the respective audio features of each channel of the multi-channel audio. The method also includes generating an output from an application-specific classifier. The first score indicates a likelihood that the multi-channel audio corresponds to a particular audio type that the particular application is configured to process.
    Type: Grant
    Filed: September 21, 2023
    Date of Patent: November 26, 2024
    Assignee: Google LLC
    Inventors: Nolan Andrew Miller, Ramin Mehran
  • Publication number: 20240013772
    Abstract: A method for multi-channel voice activity detection includes receiving a sequence of input frames characterizing streaming multi-channel audio captured by an array of microphones. Each channel of the streaming multi-channel audio includes respective audio features captured by a separate dedicated microphone. The method also includes determining, using a location fingerprint model, a location fingerprint indicating a location of a source of the multi-channel audio relative to the user device based on the respective audio features of each channel of the multi-channel audio. The method also includes generating an output from an application-specific classifier. The first score indicates a likelihood that the multi-channel audio corresponds to a particular audio type that the particular application is configured to process.
    Type: Application
    Filed: September 21, 2023
    Publication date: January 11, 2024
    Applicant: Google LLC
    Inventors: Nolan Andrew Miller, Ramin Mehran
  • Patent number: 11790888
    Abstract: A method for multi-channel voice activity detection includes receiving a sequence of input frames characterizing streaming multi-channel audio captured by an array of microphones. Each channel of the streaming multi-channel audio includes respective audio features captured by a separate dedicated microphone. The method also includes determining, using a location fingerprint model, a location fingerprint indicating a location of a source of the multi-channel audio relative to the user device based on the respective audio features of each channel of the multi-channel audio. The method also includes generating an output from an application-specific classifier. The first score indicates a likelihood that the multi-channel audio corresponds to a particular audio type that the particular application is configured to process.
    Type: Grant
    Filed: June 9, 2022
    Date of Patent: October 17, 2023
    Assignee: Google LLC
    Inventors: Nolan Andrew Miller, Ramin Mehran
  • Patent number: 11480433
    Abstract: Techniques are described for using computing devices to perform automated operations to generate mapping information using inter-connected images of a defined area, and for using the generated mapping information in further automated manners. In at least some situations, the defined area includes an interior of a multi-room building, and the generated information includes a floor map of the building, such as from an automated analysis of multiple panorama images or other images acquired at various viewing locations within the building—in at least some such situations, the generating is further performed without having detailed information about distances from the images' viewing locations to walls or other objects in the surrounding building. The generated floor map and other mapping-related information may be used in various manners, including for controlling navigation of devices (e.g., autonomous vehicles), for display on one or more client devices in corresponding graphical user interfaces, etc.
    Type: Grant
    Filed: September 17, 2021
    Date of Patent: October 25, 2022
    Assignee: Zillow, Inc.
    Inventors: Alex Colburn, Qi Shan, Ramin Mehran, Li Guan
  • Publication number: 20220310060
    Abstract: A method for multi-channel voice activity detection includes receiving a sequence of input frames characterizing streaming multi-channel audio captured by an array of microphones. Each channel of the streaming multi-channel audio includes respective audio features captured by a separate dedicated microphone. The method also includes determining, using a location fingerprint model, a location fingerprint indicating a location of a source of the multi-channel audio relative to the user device based on the respective audio features of each channel of the multi-channel audio. The method also includes generating an output from an application-specific classifier. The first score indicates a likelihood that the multi-channel audio corresponds to a particular audio type that the particular application is configured to process.
    Type: Application
    Filed: June 9, 2022
    Publication date: September 29, 2022
    Applicant: Google LLC
    Inventors: Nolan Andrew Miller, Ramin Mehran
  • Patent number: 11408738
    Abstract: Techniques are described for using computing devices to perform automated operations to generate mapping information using inter-connected images of a defined area, and for using the generated mapping information in further automated manners. In at least some situations, the defined area includes an interior of a multi-room building, and the generated information includes a floor map of the building, such as from an automated analysis of multiple panorama images or other images acquired at various viewing locations within the building—in at least some such situations, the generating is further performed without having detailed information about distances from the images' viewing locations to walls or other objects in the surrounding building. The generated floor map and other mapping-related information may be used in various manners, including for controlling navigation of devices (e.g., autonomous vehicles), for display on one or more client devices in corresponding graphical user interfaces, etc.
    Type: Grant
    Filed: September 12, 2020
    Date of Patent: August 9, 2022
    Assignee: Zillow, Inc.
    Inventors: Alex Colburn, Qi Shan, Ramin Mehran, Li Guan
  • Patent number: 11380302
    Abstract: A method for multi-channel voice activity detection includes receiving a sequence of input frames characterizing streaming multi-channel audio captured by an array of microphones. Each channel of the streaming multi-channel audio includes respective audio features captured by a separate dedicated microphone. The method also includes determining, using a location fingerprint model, a location fingerprint indicating a location of a source of the multi-channel audio relative to the user device based on the respective audio features of each channel of the multi-channel audio. The method also includes generating an output from an application-specific classifier. The first score indicates a likelihood that the multi-channel audio corresponds to a particular audio type that the particular application is configured to process.
    Type: Grant
    Filed: October 22, 2020
    Date of Patent: July 5, 2022
    Assignee: Google LLC
    Inventors: Nolan Andrew Miller, Ramin Mehran
  • Publication number: 20220130375
    Abstract: A method for multi-channel voice activity detection includes receiving a sequence of input frames characterizing streaming multi-channel audio captured by an array of microphones. Each channel of the streaming multi-channel audio includes respective audio features captured by a separate dedicated microphone. The method also includes determining, using a location fingerprint model, a location fingerprint indicating a location of a source of the multi-channel audio relative to the user device based on the respective audio features of each channel of the multi-channel audio. The method also includes generating an output from an application-specific classifier. The first score indicates a likelihood that the multi-channel audio corresponds to a particular audio type that the particular application is configured to process.
    Type: Application
    Filed: October 22, 2020
    Publication date: April 28, 2022
    Applicant: Google LLC
    Inventors: Nolan Andrew Miller, Ramin Mehran
  • Publication number: 20220003555
    Abstract: Techniques are described for using computing devices to perform automated operations to generate mapping information using inter-connected images of a defined area, and for using the generated mapping information in further automated manners. In at least some situations, the defined area includes an interior of a multi-room building, and the generated information includes a floor map of the building, such as from an automated analysis of multiple panorama images or other images acquired at various viewing locations within the building—in at least some such situations, the generating is further performed without having detailed information about distances from the images' viewing locations to walls or other objects in the surrounding building. The generated floor map and other mapping-related information may be used in various manners, including for controlling navigation of devices (e.g., autonomous vehicles), for display on one or more client devices in corresponding graphical user interfaces, etc.
    Type: Application
    Filed: September 17, 2021
    Publication date: January 6, 2022
    Inventors: Alex Colburn, Qi Shan, Ramin Mehran, Li Guan
  • Patent number: 10930262
    Abstract: A device for communicating with a remote device is disclosed, which includes a processor and a memory in communication with the processor. The memory includes executable instructions that, when executed, cause the processor to control the device to perform functions of establishing, via a communication network, a communication session with the remote device; capturing a speech spoken by a user and generating audio data representing the captured speech by the user; encoding the audio data for transmission to the remote device via the communication network; converting the audio data to text data representing the captured speech; and transmitting, during the communication session, the encoded audio data and the text data to the remote device via the communication network. The device thus can provide the text data representing the captured speech when a quality of the encoded audio signal received by the remote device is below a predetermined level.
    Type: Grant
    Filed: September 30, 2018
    Date of Patent: February 23, 2021
    Assignee: Microsoft Technology Licensing, LLC.
    Inventors: Ross G. Cutler, Sriram Srinivasan, Ramin Mehran, Karlton David Sequeira, Jayant Ajit Gupchup, Senthil K. Velayutham
  • Publication number: 20200408532
    Abstract: Techniques are described for using computing devices to perform automated operations to generate mapping information using inter-connected images of a defined area, and for using the generated mapping information in further automated manners. In at least some situations, the defined area includes an interior of a multi-room building, and the generated information includes a floor map of the building, such as from an automated analysis of multiple panorama images or other images acquired at various viewing locations within the building—in at least some such situations, the generating is further performed without having detailed information about distances from the images' viewing locations to walls or other objects in the surrounding building. The generated floor map and other mapping-related information may be used in various manners, including for controlling navigation of devices (e.g., autonomous vehicles), for display on one or more client devices in corresponding graphical user interfaces, etc.
    Type: Application
    Filed: September 12, 2020
    Publication date: December 31, 2020
    Inventors: Alex Colburn, Qi Shan, Ramin Mehran, Li Guan
  • Patent number: 10809066
    Abstract: Techniques are described for using computing devices to perform automated operations to generate mapping information using inter-connected images of a defined area, and for using the generated mapping information in further automated manners. In at least some situations, the defined area includes an interior of a multi-room building, and the generated information includes a floor map of the building, such as from an automated analysis of multiple panorama images or other images acquired at various viewing locations within the building—in at least some such situations, the generating is further performed without having detailed information about distances from the images' viewing locations to walls or other objects in the surrounding building. The generated floor map and other mapping-related information may be used in various manners, including for controlling navigation of devices (e.g., autonomous vehicles), for display on one or more client devices in corresponding graphical user interfaces, etc.
    Type: Grant
    Filed: November 14, 2018
    Date of Patent: October 20, 2020
    Assignee: Zillow Group, Inc.
    Inventors: Alex Colburn, Qi Shan, Ramin Mehran, Li Guan
  • Patent number: 10789685
    Abstract: A privacy image generation system may use a light field camera that includes an array of cameras or an RGBZ camera(s)) is used to capture images and display images according to a selected privacy mode. The privacy mode may include a blur background mode that can be automatically selected based on the meeting type, participants, location, and device type. A region of interest and/or an object(s) of interest (e.g. one or more persons in a foreground) is determined and the privacy image generation system is configured to clearly show the region/object of interest and obscure or replace the background by combining multiple images. The displayed image includes the region/object(s) of interest clearly shown (e.g. in focus) and any objects in a background of the combined image shown having a limited depth of field (e.g. blurry/not in focus) and/or blurred due to the combination of the multiple images.
    Type: Grant
    Filed: August 24, 2018
    Date of Patent: September 29, 2020
    Assignee: Microsoft Technology Licensing, LLC
    Inventors: Ross Cutler, Ramin Mehran
  • Publication number: 20200116493
    Abstract: Techniques are described for using computing devices to perform automated operations to generate mapping information using inter-connected images of a defined area, and for using the generated mapping information in further automated manners. In at least some situations, the defined area includes an interior of a multi-room building, and the generated information includes a floor map of the building, such as from an automated analysis of multiple panorama images or other images acquired at various viewing locations within the building—in at least some such situations, the generating is further performed without having detailed information about distances from the images' viewing locations to walls or other objects in the surrounding building. The generated floor map and other mapping-related information may be used in various manners, including for controlling navigation of devices (e.g., autonomous vehicles), for display on one or more client devices in corresponding graphical user interfaces, etc.
    Type: Application
    Filed: November 14, 2018
    Publication date: April 16, 2020
    Inventors: Alex Colburn, Qi Shan, Ramin Mehran, Li Guan
  • Publication number: 20190073993
    Abstract: A device is disclosed, which includes a processor and a memory in communication with the processor. The memory includes executable instructions that, when executed by the processor, cause the processor to control the device to perform functions of capturing a speech by a user; generating audio data representing the captured speech by a user; generating, based on the audio data, text data representing at least a portion of the captured speech; and transmitting, via a communication channel, the audio data and text data to the remote device. The device thus can provide the text data representing the captured speech when a quality of the audio signal received by the remote device is below a predetermined level.
    Type: Application
    Filed: October 31, 2018
    Publication date: March 7, 2019
    Applicant: MICROSOFT TECHNOLOGY LICENSING, LLC
    Inventors: Ross G. Cutler, Sriram Srinivasan, Ramin Mehran, Karlton David Sequeira, Jayant Ajit Gupchup, Senthil K. Velayutham
  • Publication number: 20190035383
    Abstract: A device for communicating with a remote device is disclosed, which includes a processor and a memory in communication with the processor. The memory includes executable instructions that, when executed, cause the processor to control the device to perform functions of establishing, via a communication network, a communication session with the remote device; capturing a speech spoken by a user and generating audio data representing the captured speech by the user; encoding the audio data for transmission to the remote device via the communication network; converting the audio data to text data representing the captured speech; and transmitting, during the communication session, the encoded audio data and the text data to the remote device via the communication network. The device thus can provide the text data representing the captured speech when a quality of the encoded audio signal received by the remote device is below a predetermined level.
    Type: Application
    Filed: September 30, 2018
    Publication date: January 31, 2019
    Applicant: MICROSOFT TECHNOLOGY LICENSING, LLC
    Inventors: Ross G. Cutler, Sriram Srinivasan, Ramin Mehran, Karlton David Sequeira, Jayant Ajit Gupchup, Senthil K. Velayutham
  • Patent number: 10181178
    Abstract: A privacy image generation system may use a light field camera that includes an array of cameras or an RGBZ camera(s)) is used to capture images and display images according to a selected privacy mode. The privacy mode may include a blur background mode and a background replacement mode and can be automatically selected based on the meeting type, participants, location, and device type. A region of interest and/or an object(s) of interest (e.g. one or more persons in a foreground) is determined and the privacy image generation system is configured to clearly show the region/object of interest and obscure or replace the background according to the selected privacy mode. The displayed image includes the region/object(s) of interest clearly shown (e.g. in focus) and any objects in a background of the combined image shown having a limited depth of field (e.g. blurry/not in focus) and/or the background replaced with another image and/or fill.
    Type: Grant
    Filed: June 30, 2017
    Date of Patent: January 15, 2019
    Assignee: Microsoft Technology Licensing, LLC
    Inventors: Ross Cutler, Ramin Mehran