Patents by Inventor Shuohuan WANG

Shuohuan WANG has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Patent number: 12585885
    Abstract: A method is provided. The method includes: obtaining a first sample dataset; inputting at least one first question text corresponding to at least one piece of first sample data into a dialog model separately to obtain at least one first answer prediction result; inputting each second question text into the dialog model to obtain a second answer prediction result output by the dialog model; inputting the second answer prediction result into a reward model to obtain a score of the second answer prediction result output by the reward model; determining a comprehensive loss based on the at least one first answer prediction result, a first answer text of each of the at least one piece of first sample data, and a score corresponding to each of at least one piece of second sample data; and adjusting at least one parameter of the dialog model based on the comprehensive loss.
    Type: Grant
    Filed: June 19, 2024
    Date of Patent: March 24, 2026
    Assignee: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
    Inventors: Yanbin Zhao, Siyu Ding, Shuohuan Wang, Yu Sun, Hao Tian, Hua Wu, Haifeng Wang
  • Patent number: 12579424
    Abstract: The present application discloses a method and an apparatus for adversarial training of a machine learning (ML) model and a medium. The method includes: obtaining input information in a training sample; extracting features of a plurality of input characters in the input information; inputting the features of the plurality of input characters to the ML model, to capture an attention weight on an input character of the plurality of input characters by an attention layer of the ML model; disturbing the attention weight captured by the attention layer, so that the ML model outputs a predicted character according to the attention weight disturbed; and training the ML model according to a difference between the predicted character and a labeled character in the training sample.
    Type: Grant
    Filed: July 7, 2021
    Date of Patent: March 17, 2026
    Assignee: BEIJING BAIDU NETCOM SCIENCE AND TECHNOLOGY CO., LTD.
    Inventors: Siyu Ding, Shuohuan Wang, Yu Sun
  • Publication number: 20260045000
    Abstract: An image generation method, an apparatus, an electronic device and a storage medium are provided. The method includes: discretizing a target text to obtain a plurality of text tokens; obtaining a resolution sequence based on an initial resolution and a target resolution, wherein the resolution sequence comprises a plurality of resolutions, and a difference between two adjacent resolutions of the plurality of resolutions is a preset increment; generating image tokens corresponding respectively to the plurality of resolutions based on the plurality of text tokens and the resolution sequence; and fusing all the image tokens to obtain a target image corresponding to the target text.
    Type: Application
    Filed: October 16, 2025
    Publication date: February 12, 2026
    Applicant: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
    Inventors: Zhenyu Zhang, Yi Song, Shuohuan Wang, Yu Sun, Hua Wu
  • Publication number: 20260044783
    Abstract: The task execution method includes: retrieving, from a storage unit, a hyperparameter of a target network layer in a target model; executing, using an operator unit, a first computational subtask in a computational task, according to the hyperparameter of the target network layer, so as to obtain a first feature output by the target network layer; executing, in response to reusing the hyperparameter of the target network layer, a second computational subtask in the computational task using the operator unit based on the first feature retrieved from the storage unit, so as to obtain a second feature output by the target network layer, where the first computational subtask and the second computational subtask are subtasks sequentially executed in the computational task; and determining a model output result of the target model using the operator unit based on the second feature retrieved from the storage unit.
    Type: Application
    Filed: October 21, 2025
    Publication date: February 12, 2026
    Inventors: Junyuan SHANG, Yilong CHEN, Zhenyu ZHANG, Yinqi YANG, Shuohuan WANG, Yu SUN
  • Publication number: 20260045086
    Abstract: The disclosure provides a method and an apparatus for processing an image, an electronic device, and a storage medium, which relates to the field of artificial intelligence technologies, and particularly to a technical field such as computer vision, deep learning, and large-scale models. The solution includes: obtaining an input content adapted to an image processing task, in which the input content includes at least one of: a first text token sequence, a first image token sequence, or an image-text fusion sequence; obtaining a joint feature representation including multimodal semantic information by performing cross-modal semantic modeling on the input content, in which the multimodal semantic information indicates a semantic correlation relationship of the input content in different modalities; and generating an output content adapted to the image processing task based on the joint feature representation.
    Type: Application
    Filed: October 16, 2025
    Publication date: February 12, 2026
    Applicant: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
    Inventors: Zhenyu Zhang, Yi Song, Xiaotian Han, Shuohuan Wang, Yu Sun, Hua Wu
  • Publication number: 20260045252
    Abstract: A training method for a multimodal speech language large model is provided. The implementation is: obtaining first response speech data generated by the multimodal speech language large model by inputting first inquiry speech data into the multimodal speech language large model; determining an inquiry text corresponding to the first inquiry speech data and a response text corresponding to the first response speech data; determining, based on the inquiry text and the response text, a first score; determining, based on speech features of the first inquiry speech data and speech features of the first response speech data, a second score, where the speech features include at least one of speech clarity, speech rate feature, timbre feature, intonation feature, and emotion feature; and adjusting, based on the first score and the second score, parameters of the multimodal speech language large model.
    Type: Application
    Filed: October 21, 2025
    Publication date: February 12, 2026
    Applicant: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
    Inventors: Shuohuan WANG, Junyuan SHANG, Zhenyu ZHANG, Yu SUN, Hua WU
  • Patent number: 12536432
    Abstract: A pre-training method of a neural network model, an electronic device, and a medium. The pre-training data is inputted to the initial neural network model, and the initial neural network model is pre-trained in the first training mode, in the first training mode, the plurality of hidden layers share one hidden layer parameter, and the loss value of the initial neural network model is obtained, if the loss value of the initial neural network model is less than a preset threshold, the initial neural network model continues to be pre-trained in the second training mode, in the second training mode, each of the plurality of hidden layers has its own hidden layer parameter.
    Type: Grant
    Filed: January 11, 2022
    Date of Patent: January 27, 2026
    Assignee: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
    Inventors: Yuxiang Lu, Jiaxiang Liu, Xuyi Chen, Shikun Feng, Shuohuan Wang, Yu Sun, Shiwei Huang, Jingzhou He
  • Publication number: 20260011045
    Abstract: Large model-based visual content generation and target large model training methods, relating to artificial intelligence fields such as deep learning, a large model, computer vision and natural language processing, are provided. A large model-based visual content generation method may include: obtaining target instruction information; inputting the target instruction information into a target large model to obtain and output corresponding target result information, where the target result information includes target visual content, the target result information is generated by the target large model according to target thinking information, and the target thinking information is thinking process information generated by the target large model for the target instruction information.
    Type: Application
    Filed: September 11, 2025
    Publication date: January 8, 2026
    Applicant: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
    Inventors: Shuohuan WANG, Zhenyu ZHANG, Junyuan SHANG, Yu SUN, Hua WU, Haifeng WANG
  • Patent number: 12373735
    Abstract: A method for pre-training a language model includes: constructing a pre-training language data set, in which the pre-training language data set comprises unsupervised language data and supervised language data; generating a hierarchical multi-template and multi-task language data set based on the pre-training language data set; and pre-training the language model based on the hierarchical multi-template and multi-task language data set.
    Type: Grant
    Filed: March 7, 2023
    Date of Patent: July 29, 2025
    Assignee: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
    Inventors: Junyuan Shang, Shuohuan Wang, Siyu Ding, Yanbin Zhao, Chao Pang, Yu Sun, Hao Tian, Hua Wu, Haifeng Wang
  • Publication number: 20250190811
    Abstract: A method for training a multimodal large model includes: obtaining first training data and second training data; obtaining an initial multimodal large model, wherein the multimodal large model comprises a backbone network and multiple codec networks corresponding to the multiple non-textual modalities; and the multiple codec networks perform encoding and decoding based on a same multimodal word list; performing a joint training on the multiple codec networks and the multimodal word list based on the data under the multiple non-textual modalities; and training the backbone network based on the multimodal sample reference data and the sample generation data under the target task in the second training data. The multiple codec networks perform the encoding and decoding based on the same multimodal word list, which reduces the difficulty and the cost of the model training.
    Type: Application
    Filed: February 25, 2025
    Publication date: June 12, 2025
    Applicant: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
    Inventors: Shuohuan Wang, Junyuan Shang, Yekun Chai, Yinqi Yang, Zhenyu Zhang, Yu Sun, Hua Wu, Haifeng Wang
  • Patent number: 12314677
    Abstract: A method and apparatus for pre-training a model, a device, a storage medium, and a program product. An embodiment of the method includes: acquiring a sample natural language text; generating N types of prompt words based on the sample natural language text, where N is a positive integer; generating sample input data based on the sample natural language text and the N types of prompt words; and training an initial language model based on the sample input data, to obtain a pre-trained language model.
    Type: Grant
    Filed: August 16, 2022
    Date of Patent: May 27, 2025
    Assignee: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
    Inventors: Junyuan Shang, Shuohuan Wang, Siyu Ding, Yanbin Zhao, Chao Pang, Yu Sun
  • Publication number: 20250094802
    Abstract: Provided is a model training method, a model reasoning method, an electronic device, and a storage medium, relating to the field of data processing, and especially to the technical fields of artificial intelligence, big data, deep learning and large models. The model training method includes: folding an initial token sequence for training a model based on a folding feature value for folding a token sequence to obtain at least a first token sequence subjected to the folding, wherein the initial token sequence represents a token sequence composed of T1 tokens, and the first token sequence has a sequence length less than that of the initial token sequence; and inputting at least the first token sequence into a preset model to train the preset model so as to obtain a target model.
    Type: Application
    Filed: December 2, 2024
    Publication date: March 20, 2025
    Inventors: Junyuan SHANG, Guoxia WANG, Yinqi YANG, Shuohuan WANG, Yu SUN
  • Publication number: 20250094534
    Abstract: A task execution method for a large model relates to fields of artificial intelligence, deep learning and large model technologies, and includes executing attention tasks in a task group to be fused using a target computing unit to obtain attention features, where the attention task corresponds to a weighted matrix to be fused, the weighted matrix to be fused is obtained by weighting a matrix to be fused using a weight; obtaining a processing result according to the attention features; determining a loss information according to the processing result; and weighting and fusing matrices to be fused using the target computing unit according to weights for the task group to be fused if the loss information converges, to obtain a fusion matrix for a target task group, where a target task in the target task group is executed by the target computing unit according to the fusion matrix.
    Type: Application
    Filed: December 4, 2024
    Publication date: March 20, 2025
    Inventors: Linhao ZHANG, Yilong CHEN, Junyuan SHANG, Yinqi YANG, Shuohuan WANG, Yu SUN
  • Publication number: 20250094806
    Abstract: Provided is a large language model training method, an electronic device and a storage medium, relating to the field of artificial intelligence technologies, and in particular, to the fields of deep learning, natural language processing and large model. The method includes: performing dimension reduction parameter fusion on a two-dimensional parameter matrix on each channel in each network layer in a first large language model, respectively, to obtain a second large language model; performing layer reduction parameter fusion on network layers in the second large language model based on a three-dimensional parameter matrix of each network layer in the second large language model to obtain a third large language model; and training the third large language model to obtain a target large language model under the condition that the target loss function determined based on the first and third large language models meets a preset first function condition.
    Type: Application
    Filed: December 3, 2024
    Publication date: March 20, 2025
    Inventors: Junyuan Shang, Yilong Chen, Zhenyu Zhang, Shuohuan Wang, Yu Sun, Hua Wu
  • Publication number: 20250094713
    Abstract: A multimodal data generation method is provided. The method includes: inputting a query data sequence into a multimodal model, to obtain a plurality of tokens in a response data sequence, where a current token is generated through the following operations: inputting the query data sequence and a current response data sequence into the multimodal model, so that the multimodal model generates the current token based on the query data sequence and the current response data sequence, in response to determining that the current token belongs to a first data modality; or inputting the query data sequence and a current response data sequence into the multimodal model, so that the multimodal model denoises an initial token sequence based on the query data sequence and the current response data sequence, to generate a result token sequence, in response to determining that the current token belongs to a second data modality.
    Type: Application
    Filed: December 3, 2024
    Publication date: March 20, 2025
    Applicant: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
    Inventors: Shuohuan WANG, Yekun CHAI, Siyu DING, Junyuan SHANG, Zhenyu ZHANG, Yu SUN, Hao TIAN, Hua WU, Haifeng WANG
  • Publication number: 20250061305
    Abstract: A training method, an inference method, a device, an apparatus, and a medium for a deep learning model are provided. A first model includes a plurality of first parameters, a second model comprises a plurality of second parameters, which is initialized to parameter values of a plurality of target parameters selected from the plurality of first parameters. The training method includes: determining a target loss for both the first model and the second model; adjusting parameter values, including: in response to determining that the target loss indicates that the parameter values of at least part of the target parameters need to be adjusted, synchronously adjusting the parameter values of the corresponding second parameters; and in response to determining that the target loss indicates that the parameter values of at least part of the second parameters need to be adjusted, synchronously adjusting the parameter values of the corresponding target parameters.
    Type: Application
    Filed: November 4, 2024
    Publication date: February 20, 2025
    Applicant: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
    Inventors: Shuohuan WANG, Junyuan SHANG, Yinqi YANG, Guoxia WANG, Linhao ZHANG, Yu SUN, Hua WU, Haifeng WANG
  • Patent number: 12223279
    Abstract: A method for generating a cross-lingual textual semantic model includes: acquiring a set of training data that includes pieces of monolingual non-parallel text and pieces of bilingual parallel text; determining a semantic vector of each piece of text in the set of training data by inputting each piece of text into an initial textual semantic model; determining a distance between semantic vectors of each two pieces of text in the set of training data based on the semantic vector of each piece of text in the set of training data; determining a gradient modification based on a parallel relationship between each two pieces of text in the set of training data and the distance between the semantic vectors of each two pieces of text in the set of training data; and acquiring a modified textual semantic model by modifying the initial textual semantic model based on the gradient modification.
    Type: Grant
    Filed: November 11, 2022
    Date of Patent: February 11, 2025
    Assignee: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
    Inventors: Yaqian Han, Shuohuan Wang, Yu Sun
  • Publication number: 20240412002
    Abstract: A method is provided. The method includes: obtaining a first sample dataset; inputting at least one first question text corresponding to at least one piece of first sample data into a dialog model separately to obtain at least one first answer prediction result; inputting each second question text into the dialog model to obtain a second answer prediction result output by the dialog model; inputting the second answer prediction result into a reward model to obtain a score of the second answer prediction result output by the reward model; determining a comprehensive loss based on the at least one first answer prediction result, a first answer text of each of the at least one piece of first sample data, and a score corresponding to each of at least one piece of second sample data; and adjusting at least one parameter of the dialog model based on the comprehensive loss.
    Type: Application
    Filed: June 19, 2024
    Publication date: December 12, 2024
    Applicant: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
    Inventors: Yanbin ZHAO, Siyu DING, Shuohuan WANG, Yu SUN, Hao TIAN, Hua WU, Haifeng WANG
  • Patent number: 12131728
    Abstract: The present application provides a method of training a natural language processing model, which relates to a field of artificial intelligence, and in particular to a field of natural language processing. A specific implementation scheme includes: performing a semantic learning for multi-tasks on an input text, so as to obtain a semantic feature for the multi-tasks, wherein the multi-tasks include a plurality of branch tasks; performing a feature learning for each branch task based on the semantic feature, so as to obtain a first output result for each branch task; calculating a loss for each branch task according to the first output result for the branch task; and adjusting a parameter of the natural language processing model according to the loss for each branch task. The present application further provides a method of processing a natural language, an electronic device, and a storage medium.
    Type: Grant
    Filed: May 31, 2022
    Date of Patent: October 29, 2024
    Assignee: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
    Inventors: Siyu Ding, Chao Pang, Shuohuan Wang, Yanbin Zhao, Junyuan Shang, Yu Sun, Shikun Feng, Hao Tian, Hua Wu, Haifeng Wang
  • Patent number: 12106052
    Abstract: The disclosure discloses a method and an apparatus for generating a semantic representation model, and a storage medium. The detailed implementation includes: performing recognition and segmentation on the original text included in an original text set to obtain knowledge units and non-knowledge units in the original text; performing knowledge unit-level disorder processing on the knowledge units and the non-knowledge units in the original text to obtain a disorder text; generating a training text set based on the character attribute of each character in the disorder text; and training an initial semantic representation model by employing the training text set to generate the semantic representation model.
    Type: Grant
    Filed: March 18, 2021
    Date of Patent: October 1, 2024
    Assignee: BEIJING BAIDU NETCOM SCIENCE AND TECHNOLOGY CO., LTD.
    Inventors: Shuohuan Wang, Siyu Ding, Yu Sun