Patents by Inventor Yannick Saillet

Yannick Saillet has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Patent number: 12717836
    Abstract: Described are techniques for a re-analysis of assignments of terms to assets. The techniques include detecting a change in a term ontology comprising a plurality of terms, and determining at least one selected from a group consisting of: a domain feature change vector (DFCV) for a domain of the term ontology affected by the change, and a term feature change vector (TFCV) for the term affected by the change. The techniques further include identifying assets for the re-analysis of the assignments of terms, wherein each of the identified assets is associated with an impact score value based on the DFCV and/or the TFCV, and performing the re-analysis of the assignments of terms for the identified assets ordered by the impact score value.
    Type: Grant
    Filed: February 1, 2023
    Date of Patent: August 25, 2026
    Assignee: International Business Machines Corporation
    Inventors: Oliver Suhre, Thomas Hampp-Bahnmueller, Peter Gerstl, Yannick Saillet, Albert Maier, Michael Baessler
  • Publication number: 20260170234
    Abstract: The present disclosure relates to a method for encoding character data using an encoding scheme, the encoding scheme being configured to represent, in accordance with a formatting scheme, an original binary value of each character of a predefined character set into a unique set of one or more code units of the encoding scheme, the method comprising for a specific token: representing metadata descriptive of the specific token with a binary value, referred to as metadata binary value; for each character in the specific token: representing the character with a binary value, referred to as character binary value; creating, in accordance with the formatting scheme, a unique set of code units including at least part of the metadata binary value and character binary value; storing the specific token by storing resulting one or more sets of code units.
    Type: Application
    Filed: February 17, 2025
    Publication date: June 18, 2026
    Inventors: Alexander Merschel, Stephan Hiller, Oliver Harm, Yannick Saillet, Stefan Renner
  • Patent number: 12639550
    Abstract: Classification of cell data includes obtaining a target dataset and an artificial intelligence (AI) model trained to identify relationship(s) between cells of a row and classify whether a focus cell of the row is erroneous based on the identified relationship(s), and applying the AI model to the target dataset to identify erroneous cell(s) thereof. The applying includes selecting a row of cells of the target dataset, inputting the selected row of cells to the AI model with an identification of a focus cell, the focus cell to be classified by the AI model, classifying the focus cell to obtain a classification of the focus cell, the classifying identifying whether the focus cell is erroneous, and outputting an indication of the classification of the focus cell.
    Type: Grant
    Filed: November 18, 2021
    Date of Patent: May 26, 2026
    Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
    Inventors: Shaikh Shahriar Quader, Omar Al-Shamali, James Miller, Yannick Saillet, Albert Maier, Remus Lazar
  • Patent number: 12579127
    Abstract: Described are techniques for detecting labels incorrectly assigned to data set fields. The data of each data set field, such as those data set fields assigned to the same label, are represented using a set of characteristics. The data set fields are then clustered into clusters based on the characteristics of the data of the data set fields. Those clusters of data set fields with a homogeneity (being assigned the same label) that exceeds a first threshold value and is below a second threshold value are identified. One or labels assigned to the data set fields of the identified clusters are identified as being suspect for incorrect assignments by having a frequency below a third threshold value (e.g., 3%), which may be user-designated. The label(s) identified as being suspect for incorrect assignment are then presented to a user for review.
    Type: Grant
    Filed: July 8, 2023
    Date of Patent: March 17, 2026
    Assignee: International Business Machines Corporation
    Inventors: Orna Raz, Yannick Saillet, Maya Zohar, Marcel Zalmanovici
  • Patent number: 12554689
    Abstract: A method, computer system, and a computer program product are provided for cleansing steps. These are in accordance with existing data quality and rules in existence for using a plurality of different transformation assets. Information is obtained about a plurality of different transformation assets and their associated data quality and rules are extracted. A plurality of possible cleansing steps to be performed are identified for the plurality of different transformation assets. An analysis is performed for the identified cleansing steps, on impact on the different transformation assets. It is then determined when more than one identified step has a similar semantics across the plurality of different transformation assets and when more than any two of them need to perform a similar step across the same dataset. The relevance of each cleansing step to be performed is then determined and a cleansing step order of performance is provided.
    Type: Grant
    Filed: June 12, 2023
    Date of Patent: February 17, 2026
    Assignee: International Business Machines Corporation
    Inventors: Alexander Lang, Albert Maier, Werner Schuetz, Sergej Schuetz, Martin Anton Oberhofer, Yannick Saillet, Mike W. Grasselt
  • Patent number: 12547907
    Abstract: An embodiment for monitoring machine learning models to detect and rectify model drift using governance. The embodiment may receive a plurality of machine learning models and register the plurality of machine learning models to a governance dashboard. The embodiment may automatically monitor the received plurality of machine learning models to identify factors used by each of the received plurality of machine learning models and generate corresponding clusters of similar machine learning models. The embodiment may automatically detect an incorrect decision made by a target machine learning model and then automatically calculate a correlation score between the target machine learning model and machine learning models within an associated corresponding cluster of similar machine learning models. The embodiment may, in response to detecting a correlation score above a threshold, automatically determine and output a cluster reinforcement recommendation.
    Type: Grant
    Filed: September 21, 2022
    Date of Patent: February 10, 2026
    Assignee: International Business Machines Corporation
    Inventors: Neerju Gupta, Namit Kabra, Yannick Saillet
  • Patent number: 12417306
    Abstract: Several aspects for optimizing unstructured document analysis comprise operating a document system, where the document system comprises a plurality of documents comprising unstructured content and a full-text index; receiving a request to identify documents comprising a type of data elements; selecting a sample out of the plurality of documents; determining data elements of the type in the sample of documents; determining an indicator context expression for the type of data elements out of the determined data elements of the type; determining a query for searching, using a search engine, the full-text index using the indicator context expression; and determining the documents in the document system being compliant to the determined query.
    Type: Grant
    Filed: December 19, 2022
    Date of Patent: September 16, 2025
    Assignee: International Business Machines Corporation
    Inventors: Thomas Hampp-Bahnmueller, Michael Baessler, Yannick Saillet
  • Patent number: 12386660
    Abstract: According to a computer-implemented method, an available amount of each of multiple computing resources is determined by machine logic over a period of time at a computing device. The machine logic also determines an expected usage of each computing resource to execute each workflow in a queue. The machine logic also determines a time of execution of each workflow in the queue based on the available amount of each of the multiple computing resources over time and the expected usage of each computing resource to execute each workflow in the queue.
    Type: Grant
    Filed: October 19, 2021
    Date of Patent: August 12, 2025
    Assignee: International Business Machines Corporation
    Inventors: Yannick Saillet, Namit Kabra
  • Patent number: 12380240
    Abstract: In an approach, a processor receives a request of a document. A processor identifies a set of datasets comprising a sensitive dataset, the set of datasets being interrelated in accordance with a relational model. A processor extracts attribute values of the document. A processor determines that a set of one or more attribute values of the extracted attribute values is in the set of datasets, the set of attribute values being values of a set of attributes. A processor determines that one or more entities of the sensitive dataset can be identified based on relations of the relational model between the set of attributes, where at least part of attribute values of the one or more entities comprises sensitive information. A processor, responsive to determining that the one or more entities can be identified, masks at least part of the set of one or more attribute values in the document.
    Type: Grant
    Filed: September 25, 2020
    Date of Patent: August 5, 2025
    Assignee: International Business Machines Corporation
    Inventors: Yannick Saillet, Albert Maier, Mike W. Grasselt, Michael Baessler, Lars Bremer
  • Patent number: 12380074
    Abstract: The present disclosure relates to a method of metadata enrichment using an enrichment comprising multiple steps. The method comprises: determining for an input data asset a metadata value descriptive of the input data asset. Characteristics of the metadata value of the input data asset may be determined. At least one informativeness score of the metadata value of the input data asset may be computed using the determined characteristics. An execution of the enrichment step may be skipped in case an input characteristic of the enrichment step is not part of the determined characteristics. In case the input characteristic of the enrichment step is part of the determined characteristics, the enrichment step may be adapted and executed or the enrichment step may be executed without adaptation. Labels resulting from the executed enrichment steps may be combined for providing one or more labels of the data asset.
    Type: Grant
    Filed: January 4, 2023
    Date of Patent: August 5, 2025
    Assignee: International Business Machines Corporation
    Inventors: Thomas Hampp-Bahnmueller, Peter Gerstl, Yannick Saillet, Michael Baessler, Albert Maier, Oliver Suhre
  • Patent number: 12271425
    Abstract: Embodiments of the present invention provide methods, computer program products, and systems. Embodiments of the present invention can condense a hierarchy in a data governance system, wherein the hierarchy comprises a root node and at least one child node comprising related sub-trees by determining, for a parent node in the hierarchy of governance system, governance terms and respective assignment relationships from a plurality of information assets, determining usage of the governance term in at least one of a plurality of governance rules, and marking a governance term of the plurality of governance terms for elimination based on the determined assignment relationships and the determined usage of the governance term in the plurality of governance rules. Embodiments of the present invention can then delete the governance term from the hierarchy if the governance term is marked for elimination.
    Type: Grant
    Filed: June 7, 2021
    Date of Patent: April 8, 2025
    Assignee: International Business Machines Corporation
    Inventors: Albert Maier, Mike W. Grasselt, Yannick Saillet, Lars Bremer, Michael Baessler
  • Patent number: 12265636
    Abstract: A database system can comprise records, each record including a set of attributes. The database system can further comprise database views, each database view representing a subset of the set of attributes. Data purpose objects indicating a subset of attributes of the set of attributes and a processing purpose can be stored. Each processing purpose can be associated with one or more entities that authorized access to the subset of attributes of the processing purpose. A request for data for a specific processing purpose and a selected view of the database views can be received. A data purpose object that indicates the specific processing purpose can be retrieved. The subset of attributes represented by the selected view can be compared with the subset of the attributes indicated in the retrieved data purpose object. Values of the subset of attributes of the selected view can be provided.
    Type: Grant
    Filed: December 8, 2021
    Date of Patent: April 1, 2025
    Assignee: International Business Machines Corporation
    Inventors: Lars Bremer, Albert Maier, Mike W. Grasselt, Yannick Saillet, Michael Baessler
  • Publication number: 20250013629
    Abstract: Described are techniques for detecting labels incorrectly assigned to data set fields. The data of each data set field, such as those data set fields assigned to the same label, are represented using a set of characteristics. The data set fields are then clustered into clusters based on the characteristics of the data of the data set fields. Those clusters of data set fields with a homogeneity (being assigned the same label) that exceeds a first threshold value and is below a second threshold value are identified. One or labels assigned to the data set fields of the identified clusters are identified as being suspect for incorrect assignments by having a frequency below a third threshold value (e.g., 3%), which may be user-designated. The label(s) identified as being suspect for incorrect assignment are then presented to a user for review.
    Type: Application
    Filed: July 8, 2023
    Publication date: January 9, 2025
    Inventors: Orna Raz, Yannick Saillet, Maya Zohar, Marcel Zalmanovici
  • Publication number: 20240411736
    Abstract: A method, computer system, and a computer program product are provided for cleansing steps. These are in accordance with existing data quality and rules in existence for using a plurality of different transformation assets. Information is obtained about a plurality of different transformation assets and their associated data quality and rules are extracted. A plurality of possible cleansing steps to be performed are identified for the plurality of different transformation assets. An analysis is performed for the identified cleansing steps, on impact on the different transformation assets. It is then determined when more than one identified step has a similar semantics across the plurality of different transformation assets and when more than any two of them need to perform a similar step across the same dataset. The relevance of each cleansing step to be performed is then determined and a cleansing step order of performance is provided.
    Type: Application
    Filed: June 12, 2023
    Publication date: December 12, 2024
    Inventors: Alexander Lang, Albert Maier, Werner Schuetz, Sergej Schuetz, Martin Anton Oberhofer, Yannick Saillet, Mike W. Grasselt
  • Patent number: 12088718
    Abstract: The exemplary embodiments disclose a method, a computer program product, and a computer system for protecting sensitive information. The exemplary embodiments may include using an inverted text index for evaluating one or more statistical measures of an index token of the inverted text index, using the one or more statistical measures for selecting a set of candidate tokens, extracting metadata from the inverted text index, associating the set of candidate tokens with respective token metadata, tokenizing at least one document resulting in one or more document tokens, comparing the one or more document tokens with the set of candidate tokens, selecting a set of document tokens to be masked, selecting at least part of the set of document tokens that comprises sensitive information according to the associated token metadata, masking the at least part of the set of document tokens, and providing one or more masked documents.
    Type: Grant
    Filed: October 19, 2020
    Date of Patent: September 10, 2024
    Assignee: International Business Machines Corporation
    Inventors: Michael Baessler, Albert Maier, Mike W. Grasselt, Yannick Saillet, Lars Bremer
  • Publication number: 20240256591
    Abstract: Described are techniques for a re-analysis of assignments of terms to assets. The techniques include detecting a change in a term ontology comprising a plurality of terms, and determining at least one selected from a group consisting of: a domain feature change vector (DFCV) for a domain of the term ontology affected by the change, and a term feature change vector (TFCV) for the term affected by the change. The techniques further include identifying assets for the re-analysis of the assignments of terms, wherein each of the identified assets is associated with an impact score value based on the DFCV and/or the TFCV, and performing the re-analysis of the assignments of terms for the identified assets ordered by the impact score value.
    Type: Application
    Filed: February 1, 2023
    Publication date: August 1, 2024
    Inventors: Oliver Suhre, Thomas Hampp-Bahnmueller, Peter Gerstl, Yannick Saillet, Albert Maier, Michael Baessler
  • Publication number: 20240242161
    Abstract: An approach is provided for computing and using a currency score. A currency score of a data element is determined as a weighted average of scores of dimensions of the data element. The dimensions include a combination of change frequency, change size, outdated value, and sources score dimensions. Relative to the data element, the change frequency dimension indicates an update frequency, the change size dimension indicates amounts of data being created, updated, and deleted per time unit, the outdated value dimension indicates a portion of values that are not semantically correct, but were semantically correct in the past, and the sources score dimension indicates a currency of input source(s). Based on the currency score, a currency of data included in the data element is evaluated. Based on the currency of the data, a remedial action is performed to improve the currency of the data.
    Type: Application
    Filed: January 18, 2023
    Publication date: July 18, 2024
    Inventors: Albert Maier, Mike W. Grasselt, Martin Anton Oberhofer, Alexander Lang, Sergej Schuetz, Yannick Saillet, Werner Schuetz
  • Patent number: 12026522
    Abstract: A database of deployed configurations, as well as attempted configurations that failed is maintained and used as reference to compare against configurations of attempted software deployments. Upon detecting a failed deployment, disclosed embodiments search the database for working configurations that most closely resemble the failed configuration, and rank the configurations based on various criteria. Disclosed embodiments may then automatically select a highest ranked working configuration, and perform an automatic upgrade of the necessary components to create a working configuration.
    Type: Grant
    Filed: April 6, 2021
    Date of Patent: July 2, 2024
    Assignee: International Business Machines Corporation
    Inventors: Krishna Kishore Bonagiri, Namit Kabra, Yannick Saillet, Mike W. Grasselt
  • Publication number: 20240202358
    Abstract: Several aspects for optimizing unstructured document analysis comprise operating a document system, where the document system comprises a plurality of documents comprising unstructured content and a full-text index; receiving a request to identify documents comprising a type of data elements; selecting a sample out of the plurality of documents; determining data elements of the type in the sample of documents; determining an indicator context expression for the type of data elements out of the determined data elements of the type; determining a query for searching, using a search engine, the full-text index using the indicator context expression; and determining the documents in the document system being compliant to the determined query.
    Type: Application
    Filed: December 19, 2022
    Publication date: June 20, 2024
    Inventors: Thomas Hampp-Bahnmueller, Michael Baessler, Yannick Saillet
  • Publication number: 20240152494
    Abstract: The present disclosure relates to a method of metadata enrichment using an enrichment comprising multiple steps. The method comprises: determining for an input data asset a metadata value descriptive of the input data asset. Characteristics of the metadata value of the input data asset may be determined. At least one informativeness score of the metadata value of the input data asset may be computed using the determined characteristics. An execution of the enrichment step may be skipped in case an input characteristic of the enrichment step is not part of the determined characteristics. In case the input characteristic of the enrichment step is part of the determined characteristics, the enrichment step may be adapted and executed or the enrichment step may be executed without adaptation. Labels resulting from the executed enrichment steps may be combined for providing one or more labels of the data asset.
    Type: Application
    Filed: January 4, 2023
    Publication date: May 9, 2024
    Inventors: Thomas Hampp-Bahnmueller, Peter Gerstl, Yannick Saillet, Michael Baessler, Albert Maier, Oliver Suhre