Patents by Inventor Scott Thomas Wisdom

Scott Thomas Wisdom has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20260229255
    Abstract: Provided are systems and methods that leverage machine learning to perform audio editing with improved precision and flexibility. Some example systems utilize a vector-based audio editing representation to condition and control a machine-learned audio editing model. This approach allows for detailed and precise control over audio edits by encoding the edits numerically and processing the encoding with a trained model. The system can handle a variety of edits, including audio generation, removal, transformation, time-shifting, and/or enhancement, thereby providing a comprehensive tool for audio manipulation.
    Type: Application
    Filed: February 5, 2026
    Publication date: August 6, 2026
    Inventors: Ron Weiss, Eduardo David Fonseca Montero, Dan Ellis, Scott Thomas Wisdom, Aren Jansen, Richard Channing Moore, III, Efthymios Tzinis, Pascal Tom Getreuer, Hakan Erdogan, John Randall Hershey, Kevin William Wilson, Manoj Plakal, Vivek Kumar
  • Publication number: 20260212879
    Abstract: A computer-implemented method is provided. The method includes receiving, by a computing device, an input audio waveform and an input textual description. The method further includes separating, by a neural network, the input audio waveform into a plurality of audio tracks. The method also includes determining, by the neural network, whether the input textual description describes an audio track of the plurality of separated audio tracks. The method additionally includes, upon a determination that the input textual description describes an audio track of the plurality of audio tracks, providing, by the computing device, the audio track corresponding to the input textual description to an interactive user interface.
    Type: Application
    Filed: December 30, 2022
    Publication date: July 23, 2026
    Inventors: Scott Thomas Wisdom, Marco Tagliasacchi, Kevin Ian Kilgour, John Randall Hershey, Beat Gfeller
  • Publication number: 20250384895
    Abstract: A computer-implemented method of applying a trained neural network for sound separation based on distance estimation is provided. The method includes receiving, by an audio input component of a computing device, an audio mixture from one or more sources. The method includes predicting, by a trained distance estimation neural network and based on the audio mixture, respective distances of the one or more sources from the audio input component. The method includes determining one or more near sounds and one or more far sounds based on the respective distances. The near sounds correspond to sources that are located within a threshold distance of the audio input component, and the far sounds correspond to sources that are not located within the threshold distance of the audio input component. The method includes providing the predicted one or more near sounds.
    Type: Application
    Filed: June 30, 2023
    Publication date: December 18, 2025
    Inventors: John Randall Hershey, Scott Thomas Wisdom, Hakan Erdogan, Malcolm Graham Slaney, Richard Francis Lyon, Kevin William Wilson, Katharine Patterson
  • Patent number: 12361964
    Abstract: Example methods include receiving training data comprising a plurality of audio clips and a plurality of textual descriptions of audio. The methods include generating a shared representation comprising a joint embedding. An audio embedding of a given audio clip is within a threshold distance of a text embedding of a textual description of the given audio clip. The methods include generating, based on the joint embedding, a conditioning vector and training, based on the conditioning vector, a neural network to: receive (i) an input audio waveform, and (ii) an input comprising one or more of an input textual description of a target audio source in the input audio waveform, or an audio sample of the target audio source, separate audio corresponding to the target audio source from the input audio waveform, and output the separated audio corresponding to the target audio source in response to the receiving of the input.
    Type: Grant
    Filed: June 24, 2022
    Date of Patent: July 15, 2025
    Assignee: Google LLC
    Inventors: Beat Gfeller, Kevin Ian Kilgour, Marco Tagliasacchi, Aren Jansen, Scott Thomas Wisdom, Qingqing Huang
  • Publication number: 20250054500
    Abstract: A system and method are disclosed. Audio input comprising the mixed audio signals is received by one or more client devices. The audio input is converted into a plurality of discrete tokens. A plurality of sound sources, each corresponding to a subset of discrete tokens of a plurality of subsets of discrete tokens, is determined using a trained machine learning model.
    Type: Application
    Filed: August 13, 2023
    Publication date: February 13, 2025
    Inventors: Hakan Erdogan, Scott Thomas Wisdom, John Hershey, Zalán Borsos, Marco Tagliasacchi, Neil Zeghidour, Xuankai Chang
  • Publication number: 20230419989
    Abstract: Example methods include receiving training data comprising a plurality of audio clips and a plurality of textual descriptions of audio. The methods include generating a shared representation comprising a joint embedding. An audio embedding of a given audio clip is within a threshold distance of a text embedding of a textual description of the given audio clip. The methods include generating, based on the joint embedding, a conditioning vector and training, based on the conditioning vector, a neural network to: receive (i) an input audio waveform, and (ii) an input comprising one or more of an input textual description of a target audio source in the input audio waveform, or an audio sample of the target audio source, separate audio corresponding to the target audio source from the input audio waveform, and output the separated audio corresponding to the target audio source in response to the receiving of the input.
    Type: Application
    Filed: June 24, 2022
    Publication date: December 28, 2023
    Inventors: Beat Gfeller, Kevin Ian Kilgour, Marco Tagliasacchi, Aren Jansen, Scott Thomas Wisdom, Qingqing Huang