Patents by Inventor Arijit Biswas
Arijit Biswas has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Publication number: 20260148742Abstract: Described herein is a method of waveform decoding, the method including the steps of: (a) receiving, by a waveform decoder, a bitstream including a finite bitrate representation of a source signal; (b) waveform decoding the finite bitrate representation of the source signal to obtain a waveform approximation of the source signal; (c) providing the waveform approximation of the source signal to a generative model that implements a probability density function, to obtain a probability distribution for a reconstructed signal of the source signal; and (d) generating the reconstructed signal of the source signal based on the probability distribution. Described are further a method and system for waveform coding and a method of training a generative model.Type: ApplicationFiled: August 26, 2025Publication date: May 28, 2026Applicants: Dolby Laboratories Licensing Corporation, DOLBY INTERNATIONAL ABInventors: Janusz Klejsa, Arijit Biswas, Lars Villemoes, Roy M. Fejgin, Cong Zhou
-
Publication number: 20260111728Abstract: Described herein is a method of generating a media bitstream to transmit parameters for updating a neural network implemented in a decoder, wherein the method includes the steps of: (a) determining at least one set of parameters for updating the neural network; (b) encoding the at least one set of parameters and media data to generate the media bitstream; and (c) transmitting the media bitstream to the decoder for updating the neural network with the at least one set of parameters. Described herein are further a method for updating a neural network implemented in a decoder, an apparatus for generating a media bitstream to transmit parameters for updating a neural network implemented in a decoder, an apparatus for updating a neural network implemented in a decoder and computer program products comprising a computer-readable storage medium with instructions adapted to cause the device to carry out said methods when executed by a device having processing capability.Type: ApplicationFiled: August 1, 2025Publication date: April 23, 2026Applicant: DOLBY INTERNATIONAL ABInventors: Christof Fersch, Arijit Biswas
-
Publication number: 20260038504Abstract: Techniques for generating a prompt for a language model to determine an action responsive to a user input, are described. In some embodiments, the system receives a user input, determines one or more application programming interfaces (APIs) configured to perform actions that are relevant to the user input and exemplars representing examples of using the APIs with respect to user inputs similar to the current user input. The system further determines device states of devices that are determined to be related to the user input and also determines other contextual information (e.g., weather information, time of day, geographic location, etc.). The system generates a prompt including the user input, the APIs, the exemplars, the device states, and the other contextual information. A language model processes the prompt to determine an action responsive to the user input and the system causes performance of the action.Type: ApplicationFiled: October 9, 2025Publication date: February 5, 2026Inventors: Hann Wang, Angeliki Metallinou, Melanie C B Gens, Arijit Biswas, Ying Shi
-
Patent number: 12518768Abstract: Described herein is a method for setting up a decoder for generating processed audio data from an audio bitstream, the decoder comprising a Generator of a Generative Adversarial Network, GAN, for processing the audio data. The method involves pre-configuring the Generator for processing of audio data with a set of parameters for the Generator, the parameters being determined by training the Generator using the full concatenated distribution. The method involves pre-configuring the decoder to determine a truncation mode for modifying the concatenated distribution and to apply the determined truncation mode to the concatenated distribution. Also described are a method of generating processed audio data from an audio bitstream using the GAN for processing audio data and a respective apparatus, as well as respective systems and computer program products.Type: GrantFiled: December 15, 2021Date of Patent: January 6, 2026Assignee: Dolby International ABInventor: Arijit Biswas
-
Publication number: 20250356873Abstract: A computer-implemented method of loss conditional training of a neural network for outputting an enhanced audio signal, the method including: randomly sampling a coefficient vector from a distribution of coefficients, wherein elements of the coefficient vector are indicative of weight coefficients corresponding to loss terms of a loss function: conditioning the neural network based on the coefficient vector; and training the conditioned neural network based on an audio training signal, wherein the training involves calculating the loss function for the audio training signal after processing by the conditioned neural network, using the weight coefficients indicated by the coefficient vector.Type: ApplicationFiled: June 7, 2023Publication date: November 20, 2025Applicant: DOLBY INTERNATIONAL ABInventor: Arijit BISWAS
-
Patent number: 12475359Abstract: Described herein is a method of determining parameters for a generative neural network for processing an audio signal, wherein the generative neural network includes an encoder stage mapping to a coded feature space and a decoder stage, each stage including a plurality of convolutional layers with one or more weight coefficients, the method comprising a plurality of cycles with sequential processes of: pruning the weight coefficients of either or both stages based on pruning control information, the pruning control information determining the number of weight coefficients that are pruned for respective convolutional layers; training the pruned generative neural network based on a set of training data; determining a loss for the trained and pruned generative neural network based on a loss function; and determining updated pruning control information based on the determined loss and a target loss. Further described are corresponding apparatus, programs, and computer-readable storage media.Type: GrantFiled: May 31, 2021Date of Patent: November 18, 2025Assignee: DOLBY INTERNATIONAL ABInventors: Arijit Biswas, Simon Plain
-
Patent number: 12462805Abstract: Techniques for generating a prompt for a language model to determine an action responsive to a user input, are described. In some embodiments, the system receives a user input, determines one or more application programming interfaces (APIs) configured to perform actions that are relevant to the user input and exemplars representing examples of using the APIs with respect to user inputs similar to the current user input. The system further determines device states of devices that are determined to be related to the user input and also determines other contextual information (e.g., weather information, time of day, geographic location, etc.). The system generates a prompt including the user input, the APIs, the exemplars, the device states, and the other contextual information. A language model processes the prompt to determine an action responsive to the user input and the system causes performance of the action.Type: GrantFiled: June 30, 2023Date of Patent: November 4, 2025Assignee: Amazon Technologies, Inc.Inventors: Hann Wang, Angeliki Metallinou, Melanie C B Gens, Arijit Biswas, Ying Shi
-
Publication number: 20250299670Abstract: Techniques for determining when speech is directed at another individual of a dialog, and storing a representation of such user-directed speech for use as context when processing subsequently-received system-directed speech are described. A system receives audio data and/or video data and determines therefrom that speech in the audio data is user-directed. Based on this, the system determine whether the speech is able to be used to perform an action by the system. If the speech is able to be used to perform an action, the system stores a natural language representation of the speech. Thereafter, when the system receives system-directed speech, the system generates a rewrite of a natural language representation of the system-directed speech based on the previously-received user-directed speech. The system then determines output data responsive to the system-directed speech using the rewritten natural language representation.Type: ApplicationFiled: June 6, 2025Publication date: September 25, 2025Inventors: Alexandros Potamianos, Arijit Biswas, Bonan Zheng, Anushree Venkatesh, Yohan Jo, Vincent Auvray, Nikolaos Malandrakis, Aaron Challenner, Xinyan Zhao, Angeliki Metallinou, David A. Jara, Jiahui Li, Ying Shi, Nikko Strom, Veerdhawal Pande
-
Patent number: 12424226Abstract: Described herein is a method of waveform decoding, the method including the steps of: (a) receiving, by a waveform decoder, a bitstream including a finite bitrate representation of a source signal; (b) waveform decoding the finite bitrate representation of the source signal to obtain a waveform approximation of the source signal; (c) providing the waveform approximation of the source signal to a generative model that implements a probability density function, to obtain a probability distribution for a reconstructed signal of the source signal; and (d) generating the reconstructed signal of the source signal based on the probability distribution. Described are further a method and system for waveform coding and a method of training a generative model.Type: GrantFiled: October 16, 2020Date of Patent: September 23, 2025Assignees: Dolby Laboratories Licensing Corporation, DOLBY INTERNATIONAL ABInventors: Janusz Klejsa, Arijit Biswas, Lars Villemoes, Roy M. Fejgin, Cong Zhou
-
Patent number: 12417775Abstract: Described herein is a method of processing an audio signal using a deep-learning-based generator, wherein the method includes the steps of: (a) inputting the audio signal into the generator for processing the audio signal; (b) mapping a time segment of the audio signal to a latent feature space representation, using an encoder stage of the generator; (c) upsampling the latent feature space representation using a decoder stage of the generator, wherein at least one layer of the decoder stage applies sinusoidal activation; and (d) obtaining, as an output from the decoder stage of the generator, a processed audio signal. Described are further a method for training said generator and respective apparatus, systems and computer program products.Type: GrantFiled: October 15, 2021Date of Patent: September 16, 2025Assignee: DOLBY INTERNATIONAL ABInventor: Arijit Biswas
-
Patent number: 12400113Abstract: Described herein is a method of generating a media bitstream to transmit parameters for updating a neural network implemented in a decoder, wherein the method includes the steps of: (a) determining at least one set of parameters for updating the neural network; (b) encoding the at least one set of parameters and media data to generate the media bitstream; and (c) transmitting the media bitstream to the decoder for updating the neural network with the at least one set of parameters. Described herein are further a method for updating a neural network implemented in a decoder, an apparatus for generating a media bitstream to transmit parameters for updating a neural network implemented in a decoder, an apparatus for updating a neural network implemented in a decoder and computer program products comprising a computer-readable storage medium with instructions adapted to cause the device to carry out said methods when executed by a device having processing capability.Type: GrantFiled: March 5, 2020Date of Patent: August 26, 2025Assignee: DOLBY INTERNATIONAL ABInventors: Christof Fersch, Arijit Biswas
-
Patent number: 12374326Abstract: Techniques for determining when speech is directed at another individual of a dialog, and storing a representation of such user-directed speech for use as context when processing subsequently-received system-directed speech are described. A system receives audio data and/or video data and determines therefrom that speech in the audio data is user-directed. Based on this, the system determine whether the speech is able to be used to perform an action by the system. If the speech is able to be used to perform an action, the system stores a natural language representation of the speech. Thereafter, when the system receives system-directed speech, the system generates a rewrite of a natural language representation of the system-directed speech based on the previously-received user-directed speech. The system then determines output data responsive to the system-directed speech using the rewritten natural language representation.Type: GrantFiled: April 28, 2023Date of Patent: July 29, 2025Assignee: Amazon Technologies, Inc.Inventors: Alexandros Potamianos, Arijit Biswas, Bonan Zheng, Anushree Venkatesh, Yohan Jo, Vincent Auvray, Nikolaos Malandrakis, Aaron Challenner, Xinyan Zhao, Angeliki Metallinou, David A Jara, Jiahui Li, Ying Shi, Nikko Strom, Veerdhawal Pande
-
Publication number: 20250210054Abstract: A method of decoding an original speech signal for hybrid adversarial-parametric speech synthesis comprising: (a) receiving quantized original linear prediction coding parameters estimated by applying linear prediction coding analysis filtering to an original speech signal and a quantized compressed representation of a residual of the original speech signal; (b) dequantizing the original linear prediction coding parameters and the compressed representation of the residual; (c) inputting the dequantized compressed representation of the residual into a decoder part of a Generator for applying adversarial mapping from the compressed residual domain to a fake (first) signal domain; (d) outputting, by the decoder part of the Generator, a fake speech signal; (e) applying linear prediction coding analysis filtering to the fake speech signal for obtaining a corresponding fake residual; (f) reconstructing the original speech signal by applying linear prediction coding cross-synthesis filtering to the fake residual andType: ApplicationFiled: March 10, 2025Publication date: June 26, 2025Applicant: DOLBY INTERNATIONAL ABInventors: Ahmed Mustafa, Arijit Biswas
-
Patent number: 12293237Abstract: Apparatuses, methods and storage medium for computing including determination of work placement on processor cores are disclosed herein. In embodiments, an apparatus may include one or more processors, devices, and/or circuitry to identify a favored core of the processor cores. The one or more processors, devices, and/or circuitry may be configured to determine whether to migrate a thread to or from the favored core. In some embodiments, the determination may be by a process executed by a driver and/or by an algorithm executed by a power control unit of the processor.Type: GrantFiled: December 19, 2023Date of Patent: May 6, 2025Assignee: Intel CorporationInventors: Guy M. Therien, Michael D. Powell, Venkatesh Ramani, Arijit Biswas, Guy G. Sotomayor
-
Publication number: 20250140281Abstract: Described herein is a computer-implemented deep-learning-based system for determining an indication of an audio quality of an input audio frame. The system comprises at least one inception block configured to receive at least one representation of an input audio frame and to map the at least one representation of the input audio frame into a feature map; and at least one fully connected layer configured to receive a feature map corresponding to the at least one representation of the input audio frame from the at least one inception block, wherein the at least one fully connected layer is configured to determine the indication of the audio quality of the input audio frame. Described are further respective methods of operating and training said system.Type: ApplicationFiled: November 30, 2021Publication date: May 1, 2025Applicant: DOLBY INTERNATIONAL ABInventors: Arijit BISWAS, Guanxin JIANG
-
Patent number: 12254889Abstract: A method of decoding an original speech signal for hybrid adversarial-parametric speech synthesis comprising: (a) receiving quantized original linear prediction coding parameters estimated by applying linear prediction coding analysis filtering to an original speech signal and a quantized compressed representation of a residual of the original speech signal; (b) dequantizing the original linear prediction coding parameters and the compressed representation of the residual; (c) inputting the dequantized compressed representation of the residual into a decoder part of a Generator for applying adversarial mapping from the compressed residual domain to a fake (first) signal domain; (d) outputting, by the decoder part of the Generator, a fake speech signal; (e) applying linear prediction coding analysis filtering to the fake speech signal for obtaining a corresponding fake residual; (f) reconstructing the original speech signal by applying linear prediction coding cross-synthesis filtering to the fake residual andType: GrantFiled: December 20, 2019Date of Patent: March 18, 2025Assignee: DOLBY INTERNATIONAL ABInventors: Ahmed Mustafa, Arijit Biswas
-
Publication number: 20250069616Abstract: Embodiments are directed to a companding method and system for reducing coding noise in an audio codec. A compression process reduces an original dynamic range of an initial audio signal through a compression process that divides the initial audio signal into a plurality of segments using a defined window shape, calculates a wideband gain in the frequency domain using a non-energy based average of frequency domain samples of the initial audio signal, and applies individual gain values to amplify segments of relatively low intensity and attenuate segments of relatively high intensity. The compressed audio signal is then expanded back to the substantially the original dynamic range that applies inverse gain values to amplify segments of relatively high intensity and attenuating segments of relatively low intensity. A QMF filterbank is used to analyze the initial audio signal to obtain a frequency domain representation.Type: ApplicationFiled: November 14, 2024Publication date: February 27, 2025Applicants: DOLBY LABORATORIES LICENSING CORPORATION, DOLBY INTERNATIONAL ABInventors: Per Hedelin, Arijit Biswas, Michael Schug, Vinay Melkote
-
Patent number: 12205581Abstract: The present disclosure relates to methods for processing a decoded audio signal and for selectively applying speech/dialog enhancement to the decoded audio signal. The present disclosure also relates to a method of operating a headset for computer-mediated reality. A method of processing a decoded audio signal comprises obtaining a measure of a cognitive load of a listener that listens to a rendering of the audio signal, determining whether speech/dialog enhancement shall be applied based on the obtained measure of the cognitive load, and performing speech/dialog enhancement based on the determination. A method of operating a headset for computer-mediated reality comprises obtaining eye-tracking data of a wearer of the headset, determining a measure of a cognitive load of the wearer of the headset based on the eye-tracking data, and outputting an indication of the cognitive load of the wearer of the headset.Type: GrantFiled: March 25, 2021Date of Patent: January 21, 2025Assignee: Dolby International ABInventor: Arijit Biswas
-
Publication number: 20250006196Abstract: Techniques for generating a prompt for a language model to determine an action responsive to a user input, are described. In some embodiments, the system receives a user input, determines one or more application programming interfaces (APIs) configured to perform actions that are relevant to the user input and exemplars representing examples of using the APIs with respect to user inputs similar to the current user input. The system further determines device states of devices that are determined to be related to the user input and also determines other contextual information (e.g., weather information, time of day, geographic location, etc.). The system generates a prompt including the user input, the APIs, the exemplars, the device states, and the other contextual information. A language model processes the prompt to determine an action responsive to the user input and the system causes performance of the action.Type: ApplicationFiled: June 30, 2023Publication date: January 2, 2025Inventors: Hann Wang, Angeliki Metallinou, Melanie C B Gens, Arijit Biswas, Ying Shi
-
Patent number: 12175994Abstract: Embodiments are directed to a companding method and system for reducing coding noise in an audio codec. A compression process reduces an original dynamic range of an initial audio signal through a compression process that divides the initial audio signal into a plurality of segments using a defined window shape, calculates a wideband gain in the frequency domain using a non-energy based average of frequency domain samples of the initial audio signal, and applies individual gain values to amplify segments of relatively low intensity and attenuate segments of relatively high intensity. The compressed audio signal is then expanded back to the substantially the original dynamic range that applies inverse gain values to amplify segments of relatively high intensity and attenuating segments of relatively low intensity. A QMF filterbank is used to analyze the initial audio signal to obtain a frequency domain representation.Type: GrantFiled: August 18, 2022Date of Patent: December 24, 2024Assignees: DOLBY INTERNATIONAL AB, DOLBY LABORATORIES LICENSING CORPORATIONInventors: Per Hedelin, Arijit Biswas, Michael Schug, Vinay Melkote