Patents by Inventor Yannick Saillet
Yannick Saillet has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Patent number: 12717836Abstract: Described are techniques for a re-analysis of assignments of terms to assets. The techniques include detecting a change in a term ontology comprising a plurality of terms, and determining at least one selected from a group consisting of: a domain feature change vector (DFCV) for a domain of the term ontology affected by the change, and a term feature change vector (TFCV) for the term affected by the change. The techniques further include identifying assets for the re-analysis of the assignments of terms, wherein each of the identified assets is associated with an impact score value based on the DFCV and/or the TFCV, and performing the re-analysis of the assignments of terms for the identified assets ordered by the impact score value.Type: GrantFiled: February 1, 2023Date of Patent: August 25, 2026Assignee: International Business Machines CorporationInventors: Oliver Suhre, Thomas Hampp-Bahnmueller, Peter Gerstl, Yannick Saillet, Albert Maier, Michael Baessler
-
Publication number: 20260170234Abstract: The present disclosure relates to a method for encoding character data using an encoding scheme, the encoding scheme being configured to represent, in accordance with a formatting scheme, an original binary value of each character of a predefined character set into a unique set of one or more code units of the encoding scheme, the method comprising for a specific token: representing metadata descriptive of the specific token with a binary value, referred to as metadata binary value; for each character in the specific token: representing the character with a binary value, referred to as character binary value; creating, in accordance with the formatting scheme, a unique set of code units including at least part of the metadata binary value and character binary value; storing the specific token by storing resulting one or more sets of code units.Type: ApplicationFiled: February 17, 2025Publication date: June 18, 2026Inventors: Alexander Merschel, Stephan Hiller, Oliver Harm, Yannick Saillet, Stefan Renner
-
Patent number: 12639550Abstract: Classification of cell data includes obtaining a target dataset and an artificial intelligence (AI) model trained to identify relationship(s) between cells of a row and classify whether a focus cell of the row is erroneous based on the identified relationship(s), and applying the AI model to the target dataset to identify erroneous cell(s) thereof. The applying includes selecting a row of cells of the target dataset, inputting the selected row of cells to the AI model with an identification of a focus cell, the focus cell to be classified by the AI model, classifying the focus cell to obtain a classification of the focus cell, the classifying identifying whether the focus cell is erroneous, and outputting an indication of the classification of the focus cell.Type: GrantFiled: November 18, 2021Date of Patent: May 26, 2026Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATIONInventors: Shaikh Shahriar Quader, Omar Al-Shamali, James Miller, Yannick Saillet, Albert Maier, Remus Lazar
-
Patent number: 12579127Abstract: Described are techniques for detecting labels incorrectly assigned to data set fields. The data of each data set field, such as those data set fields assigned to the same label, are represented using a set of characteristics. The data set fields are then clustered into clusters based on the characteristics of the data of the data set fields. Those clusters of data set fields with a homogeneity (being assigned the same label) that exceeds a first threshold value and is below a second threshold value are identified. One or labels assigned to the data set fields of the identified clusters are identified as being suspect for incorrect assignments by having a frequency below a third threshold value (e.g., 3%), which may be user-designated. The label(s) identified as being suspect for incorrect assignment are then presented to a user for review.Type: GrantFiled: July 8, 2023Date of Patent: March 17, 2026Assignee: International Business Machines CorporationInventors: Orna Raz, Yannick Saillet, Maya Zohar, Marcel Zalmanovici
-
Patent number: 12554689Abstract: A method, computer system, and a computer program product are provided for cleansing steps. These are in accordance with existing data quality and rules in existence for using a plurality of different transformation assets. Information is obtained about a plurality of different transformation assets and their associated data quality and rules are extracted. A plurality of possible cleansing steps to be performed are identified for the plurality of different transformation assets. An analysis is performed for the identified cleansing steps, on impact on the different transformation assets. It is then determined when more than one identified step has a similar semantics across the plurality of different transformation assets and when more than any two of them need to perform a similar step across the same dataset. The relevance of each cleansing step to be performed is then determined and a cleansing step order of performance is provided.Type: GrantFiled: June 12, 2023Date of Patent: February 17, 2026Assignee: International Business Machines CorporationInventors: Alexander Lang, Albert Maier, Werner Schuetz, Sergej Schuetz, Martin Anton Oberhofer, Yannick Saillet, Mike W. Grasselt
-
Patent number: 12547907Abstract: An embodiment for monitoring machine learning models to detect and rectify model drift using governance. The embodiment may receive a plurality of machine learning models and register the plurality of machine learning models to a governance dashboard. The embodiment may automatically monitor the received plurality of machine learning models to identify factors used by each of the received plurality of machine learning models and generate corresponding clusters of similar machine learning models. The embodiment may automatically detect an incorrect decision made by a target machine learning model and then automatically calculate a correlation score between the target machine learning model and machine learning models within an associated corresponding cluster of similar machine learning models. The embodiment may, in response to detecting a correlation score above a threshold, automatically determine and output a cluster reinforcement recommendation.Type: GrantFiled: September 21, 2022Date of Patent: February 10, 2026Assignee: International Business Machines CorporationInventors: Neerju Gupta, Namit Kabra, Yannick Saillet
-
Patent number: 12417306Abstract: Several aspects for optimizing unstructured document analysis comprise operating a document system, where the document system comprises a plurality of documents comprising unstructured content and a full-text index; receiving a request to identify documents comprising a type of data elements; selecting a sample out of the plurality of documents; determining data elements of the type in the sample of documents; determining an indicator context expression for the type of data elements out of the determined data elements of the type; determining a query for searching, using a search engine, the full-text index using the indicator context expression; and determining the documents in the document system being compliant to the determined query.Type: GrantFiled: December 19, 2022Date of Patent: September 16, 2025Assignee: International Business Machines CorporationInventors: Thomas Hampp-Bahnmueller, Michael Baessler, Yannick Saillet
-
Patent number: 12386660Abstract: According to a computer-implemented method, an available amount of each of multiple computing resources is determined by machine logic over a period of time at a computing device. The machine logic also determines an expected usage of each computing resource to execute each workflow in a queue. The machine logic also determines a time of execution of each workflow in the queue based on the available amount of each of the multiple computing resources over time and the expected usage of each computing resource to execute each workflow in the queue.Type: GrantFiled: October 19, 2021Date of Patent: August 12, 2025Assignee: International Business Machines CorporationInventors: Yannick Saillet, Namit Kabra
-
Patent number: 12380240Abstract: In an approach, a processor receives a request of a document. A processor identifies a set of datasets comprising a sensitive dataset, the set of datasets being interrelated in accordance with a relational model. A processor extracts attribute values of the document. A processor determines that a set of one or more attribute values of the extracted attribute values is in the set of datasets, the set of attribute values being values of a set of attributes. A processor determines that one or more entities of the sensitive dataset can be identified based on relations of the relational model between the set of attributes, where at least part of attribute values of the one or more entities comprises sensitive information. A processor, responsive to determining that the one or more entities can be identified, masks at least part of the set of one or more attribute values in the document.Type: GrantFiled: September 25, 2020Date of Patent: August 5, 2025Assignee: International Business Machines CorporationInventors: Yannick Saillet, Albert Maier, Mike W. Grasselt, Michael Baessler, Lars Bremer
-
Patent number: 12380074Abstract: The present disclosure relates to a method of metadata enrichment using an enrichment comprising multiple steps. The method comprises: determining for an input data asset a metadata value descriptive of the input data asset. Characteristics of the metadata value of the input data asset may be determined. At least one informativeness score of the metadata value of the input data asset may be computed using the determined characteristics. An execution of the enrichment step may be skipped in case an input characteristic of the enrichment step is not part of the determined characteristics. In case the input characteristic of the enrichment step is part of the determined characteristics, the enrichment step may be adapted and executed or the enrichment step may be executed without adaptation. Labels resulting from the executed enrichment steps may be combined for providing one or more labels of the data asset.Type: GrantFiled: January 4, 2023Date of Patent: August 5, 2025Assignee: International Business Machines CorporationInventors: Thomas Hampp-Bahnmueller, Peter Gerstl, Yannick Saillet, Michael Baessler, Albert Maier, Oliver Suhre
-
Patent number: 12271425Abstract: Embodiments of the present invention provide methods, computer program products, and systems. Embodiments of the present invention can condense a hierarchy in a data governance system, wherein the hierarchy comprises a root node and at least one child node comprising related sub-trees by determining, for a parent node in the hierarchy of governance system, governance terms and respective assignment relationships from a plurality of information assets, determining usage of the governance term in at least one of a plurality of governance rules, and marking a governance term of the plurality of governance terms for elimination based on the determined assignment relationships and the determined usage of the governance term in the plurality of governance rules. Embodiments of the present invention can then delete the governance term from the hierarchy if the governance term is marked for elimination.Type: GrantFiled: June 7, 2021Date of Patent: April 8, 2025Assignee: International Business Machines CorporationInventors: Albert Maier, Mike W. Grasselt, Yannick Saillet, Lars Bremer, Michael Baessler
-
Patent number: 12265636Abstract: A database system can comprise records, each record including a set of attributes. The database system can further comprise database views, each database view representing a subset of the set of attributes. Data purpose objects indicating a subset of attributes of the set of attributes and a processing purpose can be stored. Each processing purpose can be associated with one or more entities that authorized access to the subset of attributes of the processing purpose. A request for data for a specific processing purpose and a selected view of the database views can be received. A data purpose object that indicates the specific processing purpose can be retrieved. The subset of attributes represented by the selected view can be compared with the subset of the attributes indicated in the retrieved data purpose object. Values of the subset of attributes of the selected view can be provided.Type: GrantFiled: December 8, 2021Date of Patent: April 1, 2025Assignee: International Business Machines CorporationInventors: Lars Bremer, Albert Maier, Mike W. Grasselt, Yannick Saillet, Michael Baessler
-
Publication number: 20250013629Abstract: Described are techniques for detecting labels incorrectly assigned to data set fields. The data of each data set field, such as those data set fields assigned to the same label, are represented using a set of characteristics. The data set fields are then clustered into clusters based on the characteristics of the data of the data set fields. Those clusters of data set fields with a homogeneity (being assigned the same label) that exceeds a first threshold value and is below a second threshold value are identified. One or labels assigned to the data set fields of the identified clusters are identified as being suspect for incorrect assignments by having a frequency below a third threshold value (e.g., 3%), which may be user-designated. The label(s) identified as being suspect for incorrect assignment are then presented to a user for review.Type: ApplicationFiled: July 8, 2023Publication date: January 9, 2025Inventors: Orna Raz, Yannick Saillet, Maya Zohar, Marcel Zalmanovici
-
Publication number: 20240411736Abstract: A method, computer system, and a computer program product are provided for cleansing steps. These are in accordance with existing data quality and rules in existence for using a plurality of different transformation assets. Information is obtained about a plurality of different transformation assets and their associated data quality and rules are extracted. A plurality of possible cleansing steps to be performed are identified for the plurality of different transformation assets. An analysis is performed for the identified cleansing steps, on impact on the different transformation assets. It is then determined when more than one identified step has a similar semantics across the plurality of different transformation assets and when more than any two of them need to perform a similar step across the same dataset. The relevance of each cleansing step to be performed is then determined and a cleansing step order of performance is provided.Type: ApplicationFiled: June 12, 2023Publication date: December 12, 2024Inventors: Alexander Lang, Albert Maier, Werner Schuetz, Sergej Schuetz, Martin Anton Oberhofer, Yannick Saillet, Mike W. Grasselt
-
Patent number: 12088718Abstract: The exemplary embodiments disclose a method, a computer program product, and a computer system for protecting sensitive information. The exemplary embodiments may include using an inverted text index for evaluating one or more statistical measures of an index token of the inverted text index, using the one or more statistical measures for selecting a set of candidate tokens, extracting metadata from the inverted text index, associating the set of candidate tokens with respective token metadata, tokenizing at least one document resulting in one or more document tokens, comparing the one or more document tokens with the set of candidate tokens, selecting a set of document tokens to be masked, selecting at least part of the set of document tokens that comprises sensitive information according to the associated token metadata, masking the at least part of the set of document tokens, and providing one or more masked documents.Type: GrantFiled: October 19, 2020Date of Patent: September 10, 2024Assignee: International Business Machines CorporationInventors: Michael Baessler, Albert Maier, Mike W. Grasselt, Yannick Saillet, Lars Bremer
-
Publication number: 20240256591Abstract: Described are techniques for a re-analysis of assignments of terms to assets. The techniques include detecting a change in a term ontology comprising a plurality of terms, and determining at least one selected from a group consisting of: a domain feature change vector (DFCV) for a domain of the term ontology affected by the change, and a term feature change vector (TFCV) for the term affected by the change. The techniques further include identifying assets for the re-analysis of the assignments of terms, wherein each of the identified assets is associated with an impact score value based on the DFCV and/or the TFCV, and performing the re-analysis of the assignments of terms for the identified assets ordered by the impact score value.Type: ApplicationFiled: February 1, 2023Publication date: August 1, 2024Inventors: Oliver Suhre, Thomas Hampp-Bahnmueller, Peter Gerstl, Yannick Saillet, Albert Maier, Michael Baessler
-
Publication number: 20240242161Abstract: An approach is provided for computing and using a currency score. A currency score of a data element is determined as a weighted average of scores of dimensions of the data element. The dimensions include a combination of change frequency, change size, outdated value, and sources score dimensions. Relative to the data element, the change frequency dimension indicates an update frequency, the change size dimension indicates amounts of data being created, updated, and deleted per time unit, the outdated value dimension indicates a portion of values that are not semantically correct, but were semantically correct in the past, and the sources score dimension indicates a currency of input source(s). Based on the currency score, a currency of data included in the data element is evaluated. Based on the currency of the data, a remedial action is performed to improve the currency of the data.Type: ApplicationFiled: January 18, 2023Publication date: July 18, 2024Inventors: Albert Maier, Mike W. Grasselt, Martin Anton Oberhofer, Alexander Lang, Sergej Schuetz, Yannick Saillet, Werner Schuetz
-
Patent number: 12026522Abstract: A database of deployed configurations, as well as attempted configurations that failed is maintained and used as reference to compare against configurations of attempted software deployments. Upon detecting a failed deployment, disclosed embodiments search the database for working configurations that most closely resemble the failed configuration, and rank the configurations based on various criteria. Disclosed embodiments may then automatically select a highest ranked working configuration, and perform an automatic upgrade of the necessary components to create a working configuration.Type: GrantFiled: April 6, 2021Date of Patent: July 2, 2024Assignee: International Business Machines CorporationInventors: Krishna Kishore Bonagiri, Namit Kabra, Yannick Saillet, Mike W. Grasselt
-
Publication number: 20240202358Abstract: Several aspects for optimizing unstructured document analysis comprise operating a document system, where the document system comprises a plurality of documents comprising unstructured content and a full-text index; receiving a request to identify documents comprising a type of data elements; selecting a sample out of the plurality of documents; determining data elements of the type in the sample of documents; determining an indicator context expression for the type of data elements out of the determined data elements of the type; determining a query for searching, using a search engine, the full-text index using the indicator context expression; and determining the documents in the document system being compliant to the determined query.Type: ApplicationFiled: December 19, 2022Publication date: June 20, 2024Inventors: Thomas Hampp-Bahnmueller, Michael Baessler, Yannick Saillet
-
Publication number: 20240152494Abstract: The present disclosure relates to a method of metadata enrichment using an enrichment comprising multiple steps. The method comprises: determining for an input data asset a metadata value descriptive of the input data asset. Characteristics of the metadata value of the input data asset may be determined. At least one informativeness score of the metadata value of the input data asset may be computed using the determined characteristics. An execution of the enrichment step may be skipped in case an input characteristic of the enrichment step is not part of the determined characteristics. In case the input characteristic of the enrichment step is part of the determined characteristics, the enrichment step may be adapted and executed or the enrichment step may be executed without adaptation. Labels resulting from the executed enrichment steps may be combined for providing one or more labels of the data asset.Type: ApplicationFiled: January 4, 2023Publication date: May 9, 2024Inventors: Thomas Hampp-Bahnmueller, Peter Gerstl, Yannick Saillet, Michael Baessler, Albert Maier, Oliver Suhre