Patents Examined by Eunice Lee
-
Patent number: 12676160Abstract: Embodiments of this application provide a sound signal processing method and an electronic device, which can reduce an interference sound signal in a sound during video recording and improve quality of a sound signal during the video recording. The method is applied to an electronic device. The method includes: acquiring, by the electronic device, a first sound signal; the first sound signal being a sound signal during video recording; processing, by the electronic device, the first sound signal to obtain a second sound signal; and outputting, by the electronic device, the second sound signal when playing back a recorded video file. Energy of a sound signal in the second sound signal in a non-target orientation is lower than energy of a sound signal in the first sound signal in the non-target orientation. The non-target orientation is an orientation outside a field of view range of a camera during video recording.Type: GrantFiled: May 26, 2022Date of Patent: July 7, 2026Assignee: BEIJING HONOR DEVICE CO., LTD.Inventors: Jianyong Xuan, Zhenyi Liu, Haikuan Gao
-
Patent number: 12664985Abstract: Systems and methods configured to generate object vestiges based on audio information conveying natural speech are disclosed. Exemplary implementations may: obtain audio information captured by a client computing platform; provide the audio information to a trained intent recognition model, wherein the trained intent recognition model is trained to determine one or more semantic objects based on the audio information and subsequently determine object vestiges based on the one or more semantic objects, wherein the semantic objects indicate entities and have intent types, wherein the intent types include a content type and a directive type, wherein the entities are spoken by the participants or referred to by the participants; obtain, from the trained intent recognition model, the object vestiges; and store the object vestiges to an electronic record of the subject.Type: GrantFiled: May 29, 2024Date of Patent: June 23, 2026Assignee: Suki AI, Inc.Inventors: Matt Pallakoff, Akshay Kore, Luis Daniel Mosquera, Bahador Saket, Gaurav Trivedi
-
Patent number: 12657404Abstract: A method of enabling a language model trained in a first language to support a second language includes extending an existing vocabulary of the language model to include additional tokens for text in the second language. The method also includes initializing the additional tokens for text in the second language based on subtokens of tokens for text in the first language from the existing vocabulary. The method further includes training the language model using a mixed language dataset that includes a first language corpus and a second language corpus. In addition, the method includes performing instruction tuning using a dataset that includes (i) instruction and response pairs involving the first language and (ii) instruction and response pairs involving the second language.Type: GrantFiled: December 29, 2023Date of Patent: June 16, 2026Assignee: Samsung Electronics Co., Ltd.Inventors: Hai Wang, Zheng Tang, Vijay Srinivasan, Hyuk Joon Kwon, Vikas Yadav, Feixuan Wang, Hongxia Jin
-
Patent number: 12620405Abstract: Example signal processing methods and example electronic devices are disclosed. One example method is applied to an electronic device, where the electronic device includes a microphone array and a camera. The example method includes performing sound source localization on a first audio signal obtained by using the microphone array, to obtain sound source direction information. A first video obtained by using the camera is processed to obtain user direction information. A target sound source direction is determined based on the sound source direction information and the user direction information. A user lip video is obtained in the target sound source direction by using the camera. A second audio signal is obtained by using the microphone array. A third audio signal is obtained based on the second audio signal and the user lip video by using a voice quality enhancement model.Type: GrantFiled: September 17, 2021Date of Patent: May 5, 2026Assignee: Huawei Technologies Co., LTD.Inventors: Guangzhao Bao, Liwen Chen, Lei Huang
-
Patent number: 12614557Abstract: A system to improve pilot voice communication includes: a microphone capturing pilot speech during operational use of an aircraft; and an audio subsystem that stores recordings of the captured pilot speech during different periods of the operational use. A pilot recording selection graphic user interface permits selection by the pilot of recordings made during a low stress period of his/her operational use of the aircraft, and one or more recordings made during a high stress period. A training algorithm analyzes characteristics of the selected recordings made during the low stress period to set a baseline. The training algorithm subsequently analyzes real-time pilot speech to ascertain when its characteristics are increased by a threshold amount over the baseline, which is classified as speech made under stress. The audio system alters and improves the analyzed real-time speech made under stress by converting it to normal-sounding speech, minimizing propagation of stressed speech.Type: GrantFiled: November 1, 2023Date of Patent: April 28, 2026Assignee: TTM Technologies, Inc.Inventor: Carlton Rubio
-
Patent number: 12603078Abstract: Methods, systems, and computer program products for generating speech data using artificial intelligence techniques are provided herein. A computer-implemented method includes implementing one or more artificial intelligence techniques in connection with one or more speech synthesis tasks; generating, in multiple sequential portions, at least one sequence of data, comprising one or more of phonetic data and prosodic data, by processing at least one previously generated sequence of data using the one or more artificial intelligence techniques; and generating speech data corresponding to at least a portion of the sequence of data by processing the at least a portion of the sequence of data using at least one artificial intelligence-based speech synthesis model.Type: GrantFiled: August 15, 2023Date of Patent: April 14, 2026Assignee: International Business Machines CorporationInventors: Zvi Kons, Ron Hoory, Vyacheslav Shechtman, Avihu Dekel
-
Patent number: 12597365Abstract: Methods, apparatus, systems, and articles of manufacture to translation between sign language and spoken language are disclosed. An example apparatus includes processor circuitry to at least one of instantiate or execute machine readable instructions to identify a plurality of candidate signs across different frames in video; associate a respective gloss to respective ones of the candidate signs; associate a respective confidence score with the respective glosses; identify overlapping frames of the candidate signs; select one or more of the candidate signs as performed signs based on the respective confidence scores and overlapping frames; and convert the performed signs to audio data.Type: GrantFiled: November 22, 2022Date of Patent: April 7, 2026Assignee: Sorenson IP Holdings, LLCInventors: Mariam Rahmani, Adam Munder, Marina Lovell, Abolfazl Zargari Khuzani, John K. Hines, Katalin Bartfai-Walcott, Naveen Kulkarni, Shashank Bujimalla Venkata Sesha, Abolfazl Ravanshad
-
Patent number: 12585876Abstract: A method of training a part-of-speech (POS) tagging model includes: separating an input sentence into units of syllables to generate an input sequence; encoding, using at least one encoder included in a part-of-speech (POS) tagging model, the input sequence; generating, based on the encoded input sequence and using a first discriminator included in the POS tagging model, a POS tagging result; and generating, based on the encoded input sequence and using a second discriminator included in the POS tagging model, a spacing result.Type: GrantFiled: June 5, 2023Date of Patent: March 24, 2026Assignees: Hyundai Motor Company, Kia CorporationInventors: Cheoneum Park, Juae Kim, Cheongjae Lee, Soo Jong Do
-
Patent number: 12579385Abstract: A method of facilitating consumption of online content includes receiving source text for a source article to be translated, the source text being in a source language. The source language for the source text and the target language to which the source text is to be translated are each identified. The source text, the source language, and the target language are each provided to a machine translation model which automatically generates translated text in the target language from the source text. The translated text is provided as input to a generative language model which generates summary text in the target language from the translated text. The summary text is provided to a text-to-speech model which generates summary audio from the summary text. The summary text and summary audio are then sent to a user interface via which the summary text is displayed, and playback of the summary audio is enabled.Type: GrantFiled: November 24, 2023Date of Patent: March 17, 2026Assignee: Microsoft Technology Licensing, LLCInventors: Shveta Verma, Deepak Achuthan Menon, Amit Dangwal, Prakash Arjunan, Rupeshkumar Rasiklal Mehta, Kishor Chamua, Arijit Mukherjee, Shubham Bansal
-
Patent number: 12566928Abstract: The present disclosure relates to methods and systems that generate a confidence score for the generated large language model (LLM) output. The methods and systems use the text of the input provided to the LLM and the text from the generated LLM output to produce a feature vector that encodes a readability of the text from the input and the text of the LLM output. The feature vector is used to determine a corresponding confidence score for the generated LLM output. The confidence score is used to evaluate a quality of the generated LLM output.Type: GrantFiled: April 27, 2023Date of Patent: March 3, 2026Assignee: Microsoft Technology Licensing, LLCInventors: Shima Imani, Harsh Shrivastava
-
Patent number: 12554939Abstract: A method, a structure, and a computer system for diverse natural language generation. The exemplary embodiments may include training a data-to-text neural network (D2T NN) and training a text-to-text neural network (T2T NN), wherein the D2T NN and the T2T NN have identical transformer architectures, and wherein the training of the D2T NN and the T2T NN are indexed by time. The exemplary embodiments may further include interleaving weights between the D2T NN and the T2T NN, as well as generating a sentence based on the interleaved D2T NN and T2T NN.Type: GrantFiled: June 22, 2023Date of Patent: February 17, 2026Assignee: International Business Machines CorporationInventors: Aaron K. Baughman, Chandankumar Johakhim Patel, Mauro Marzorati, Jeremy R. Fox
-
Patent number: 12555576Abstract: A method of controlling an artificial intelligence device may include receiving an operation command at a plurality of artificial intelligence devices; determining a first artificial intelligence device closest to a point-of-origin of an operation command based on the operation command being received at the plurality of artificial intelligence devices; outputting a response corresponding to the operation command through the first artificial intelligence device; determining a second artificial intelligence device that will perform an operation corresponding to the operation command, and transmitting a control command corresponding to the operation command to the second artificial intelligence device; and performing, by the second artificial intelligence device, an operation corresponding to the operation command based on the control command.Type: GrantFiled: November 4, 2022Date of Patent: February 17, 2026Assignee: LG ELECTRONICS INC.Inventors: Yuyong Jeon, Heewan Park, Donghoon Yi
-
Patent number: 12525221Abstract: In an embodiment a system includes a training data preparation device configured to obtain a speech recognition rate of speech data for training using a target speech recognition model, a recognition rate prediction model configured to estimate an expected recognition rate of the target speech recognition model for clean speech data in which noise is removed from the speech data for training and a speech preprocessing model configured to preprocess the speech data for training to obtain the clean speech data and to update the speech preprocessing model based on a recognition rate loss corresponding to a difference between the expected recognition rate and a maximum recognition rate.Type: GrantFiled: April 4, 2023Date of Patent: January 13, 2026Assignees: Hyundai Motor Company, Kia CorporationInventor: Yong Hyeok Lee
-
Patent number: 12518775Abstract: A method and an apparatus for evaluating voice quality are provided. In the method, a playback of a standard audio is recorded to obtain a to-be-evaluated signal. Then a first power spectrum of the to-be-evaluated signal on a critical frequency band is determined to obtain a first spectrogram. Then a second power spectrum of a reference signal corresponding to the standard audio on a critical frequency band is determined to obtain a second spectrogram, where the reference signal is a sampled signal corresponding to the standard audio. Then an image similarity between the first spectrogram and the second spectrogram is determined to obtain a voice quality score of the to-be-evaluated signal.Type: GrantFiled: September 22, 2021Date of Patent: January 6, 2026Assignee: Tencent Music Entertainment Technology (Shenzhen) Co., Ltd.Inventor: Chaopeng Zhang
-
Patent number: 12518764Abstract: Broadly speaking, the present techniques generally relate to a system, computer-implemented method and apparatus for training a machine learning, ML, model to perform sound enhancement for a target user in real-time, and to a method and apparatus for using the trained ML model to perform sound enhancement of audio signals in real-time. Advantageously, the present techniques are suitable for implementation on resource-constrained devices that capture audio signals, such as smartphones and Internet of Things devices.Type: GrantFiled: May 23, 2023Date of Patent: January 6, 2026Assignee: Samsung Electronics Co., Ltd.Inventors: Alberto Gil Ramos, Carlos Purves, Abhinav Mehrotra, Sourav Bhattacharya, Ravichander Vipperla, Syed Samin Ishtiaq, Nicholas Donald Atkins Lane
-
Patent number: 12511095Abstract: A system and method for providing targeted crowd-based acoustic-prosodic and linguistic accommodation includes simultaneously detecting, with a multimedia device, individual facial images of a plurality of individuals and acoustic sounds from the plurality of individuals. Message data representative of an audible message supplied to the plurality of individuals is supplied to the processing system, the message data. In the processing system: a plurality of facial features and a plurality of acoustic-related features are extracted from the multimedia data; the facial features and the acoustic-related features are processed to determine an aggregate emotional state of the plurality of individuals; one or more acoustic-prosodic features of the audible message and/or a content of the audible message are selectively manipulated based on the determined aggregate emotional state of the plurality of individuals to generate an updated audible message; and the updated audible message is output.Type: GrantFiled: December 6, 2022Date of Patent: December 30, 2025Assignee: HONEYWELL INTERNATIONAL INC.Inventors: Nichola Lubold, Tor Finseth
-
Patent number: 12505310Abstract: The disclosure relates to systems and methods of improved attention for neural networks. A system may access a first array of entries and a second array of entries, wherein the first array represents a first word embedding generated from one or more first characters of an input and the second array represents a second word embedding from the one or more second characters of the input. A system may generate at least a first difference matrix based on the first array and at least a second difference matrix based on the second array. A system may determine a difference value based on vector based at least in part on the difference value. A system may provide the context vector to subsequent layers of a neural network.Type: GrantFiled: April 26, 2023Date of Patent: December 23, 2025Assignee: THE BANK OF NEW YORK MELLONInventors: Christopher Policastro, Yikai Feng, Abhinav Prasad
-
Patent number: 12488197Abstract: Systems and methods for operating an artificial intelligence and machine learning model to generate annotations of data and to generate training datasets is disclosed. One disclosed system includes one or more processors configured to: assign a task to a machine learning model; receive an output from the machine learning model associated with the task; compare the output to task data associated with a user performing the assigned task; and when there is a difference between the output of the machine learning model and the task data, generate an annotation.Type: GrantFiled: April 24, 2023Date of Patent: December 2, 2025Assignee: Wells Fargo Bank, N.A.Inventors: Gabriela Garcia, Weicheng Liu, Hope Ann Staroselsky
-
Patent number: 12488795Abstract: The technology disclosed herein enables provision of sales guidance to an agent on a real-time communication session based on background sound identified during the communication session. In a particular embodiment, a method includes receiving audio from a first endpoint operated by a first user. The audio is received over a real-time communication session established between the first endpoint and a second endpoint operated by an agent of a contact center. The method further includes identifying sound other than a voice of the first user from the audio and determining a characteristic of the first user indicated by the sound. During the communication session, the method includes providing sales guidance to the agent based on the characteristic.Type: GrantFiled: November 29, 2022Date of Patent: December 2, 2025Assignee: Avaya Management L.P.Inventors: Rusty Gerald Nelson, Paul Roller Michaelis, Kevin Archer, Gregory Paul Schin
-
Patent number: 12481691Abstract: Systems and methods for data narration are provided. One aspect of the systems and methods includes obtaining a dataset including a plurality of data elements, wherein each of the data elements includes a plurality of attributes; clustering the plurality of data elements to obtain a plurality of data segments; extracting a plurality of facts from the dataset based on the plurality of data segments; generating a graph including a plurality of nodes corresponding to the plurality of facts, respectively; computing an ordering of the plurality of facts based on the graph; and generating a description of the dataset based on the ordering of the plurality of facts.Type: GrantFiled: January 26, 2023Date of Patent: November 25, 2025Assignee: ADOBE INC.Inventors: Vibhor Porwal, Aaron Jerry Ninan, Adit Akarsh, Aryan Yadav, Nischay ., Ramasuri Narayanam, Iftikhar Ahamath Burhanuddin