What technology area does this patent fall under?

Primary CPC classification G06F16/215. Mapped technology areas include Physics.

When was this patent published?

Publication date Tue Dec 02 2025 00:00:00 GMT+0000 (Coordinated Universal Time) (B2). Legal status and post-grant events are not shown on this page.

What related patents are in patentsdb?

We list 3 related publications on this page (citations in our corpus or others sharing the same primary CPC).

Automatically improving data annotations by processing annotation properties and user feedback

US12487976B2 · US · B2

Patent metadata
Field	Value
Publication number	US-12487976-B2
Application number	US-202117494987-A
Country	US
Kind code	B2
Filing date	Oct 6, 2021
Priority date	Oct 6, 2021
Publication date	Dec 2, 2025
Grant date	Dec 2, 2025

How to read this patent

A practical reading order for non-experts. Skip the full description unless you need deep technical detail.

Title
What the patent document calls the invention.
Abstract
A short plain-language summary of the technical disclosure.
Assignees and inventors
Who owns or filed the patent and who is credited as inventor.
Key dates
Filing, priority, publication, and grant dates set the timeline.
First independent claim
The legal scope of protection — read this for what is actually claimed.
CPC / IPC classifications
Technology tags used to group this patent with similar filings.
Citations and related patents
Prior art links and similar publications in this corpus.

Abstract

Official abstract text for this publication.

Methods, systems, and computer program products for automatically improving data annotations by processing annotation properties and user feedback are provided herein. A computer-implemented method includes obtaining data annotation pairs, each comprising an input data annotation in a first format and a corresponding output data annotation in a second format; determining, within at least a portion of the data annotation pairs, one or more non-diffs; identifying, across the at least a portion of data annotation pairs, data annotation properties associated with multiple intents by processing the non-diffs using property-related rules; modifying at least a portion of the data annotation pairs based on the identified data annotation properties; outputting the modified data annotation pairs to at least one user; and generating a final collection of data annotation pairs by processing at least a portion of the modified data annotation pairs and user feedback received in response to the outputting.

First claim

Opening claim text (preview).

What is claimed is: 1 . A computer-implemented method comprising: obtaining a set of data annotation pairs, wherein each of the data annotation pairs comprises an input data annotation in a first format and a corresponding output data annotation in a second format; determining, within at least a portion of the data annotation pairs, one or more non-diffs; determining, across the at least a portion of the data annotation pairs, one or more data annotation properties associated with multiple intents by processing at least a portion of the one or more non-diffs using one or more regular expression learning-based clustering algorithms to group instances of the one or more non-diffs within the at least a portion of the data annotation pairs on a basis of at least one of (i) one or more repeating characters within the one or more non-diffs, (ii) non-diff positioning, and (iii) one or more matching words within the one or more non-diffs; modifying at least a portion of the data annotation pairs based at least in part on the one or more identified data annotation properties; outputting the modified data annotation pairs to at least one user; and generating a final collection of data annotation pairs by processing at least a portion of the modified data annotation pairs and user feedback received in response to the outputting of the modified data annotation pairs; wherein the method is carried out by at least one computing device. 2 . The computer-implemented method of claim 1 , wherein determining the one or more non-diffs comprises processing the at least a portion of the data annotation pairs for presence of similar length non-diffs in respective input data annotations and corresponding output data annotations. 3 . The computer-implemented method of claim 1 , wherein determining the one or more non-diffs comprises processing the at least a portion of the data annotation pairs for presence of identical position placement of at least one non-diff in respective input data annotations and corresponding output data annotations. 4 . The computer-implemented method of claim 1 , wherein determining the one or more non-diffs comprises processing the at least a portion of the data annotation pairs for presence of at least one non-diff of a same token type in respective input data annotations and corresponding output data annotations. 5 . The computer-implemented method of claim 1 , wherein determining the one or more non-diffs comprises processing the at least a portion of the data annotation pairs for presence of one or more repeating characters within at least one non-diff in respective input data annotations and corresponding output data annotations. 6 . The computer-implemented method of claim 1 , wherein the user feedback comprises at least one of acceptance of at least a portion of the one or more identified data annotation properties and rejection of at least a portion of the one or more identified data annotation properties. 7 . The computer-implemented method of claim 1 , wherein modifying the at least a portion of the data annotation pairs comprises generating one or more new data annotation pairs based at least in part on the one or more identified data annotation properties and the obtained set of data annotation pairs. 8 . The computer-implemented method of claim 1 , wherein modifying the at least a portion of the data annotation pairs comprises updating one or more of the obtained set of data annotation pairs based at least in part on the one or more identified data annotation properties. 9 . The computer-implemented method of claim 1 , further comprising: computing quality scores for the obtained set of data annotation pairs, wherein the quality scores are based at least in part on extent of multiple intents associated with the obtained set of data annotation pairs. 10 . The computer-implemented method of claim 9 , further comprising: computing quality scores for the final collection of data annotation pairs, wherein the quality scores are based at least in part on extent of multiple intents associated with the final collection of data annotation pairs; and comparing the quality scores for the final collection of data annotation pairs to the quality scores for the obtained set of data annotation pairs. 11 . The computer-implemented method of claim 1 , wherein software implementing the method is provided as a service in a cloud environment. 12 . A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computing device to cause the computing device to: obtain a set of data annotation pairs, wherein each of the data annotation pairs comprises an input data annotation in a first format and a corresponding output data annotation in a second format; determine, within at least a portion of the data annotation pairs, one or more non-diffs; determine, across the at least a portion of the data annotation pairs, one or more data annotation properties associated with multiple intents by processing at least a portion of the one or more non-diffs using one or more regular expression learning-based clustering algorithms to group instances of the one or more non-diffs within the at least a portion of the data annotation pairs on a basis of at least one of (i) one or more repeating characters within the one or more non-diffs, (ii) non-diff positioning, and (iii) one or more matching words within the one or more non-diffs; modify at least a portion of the data annotation pairs based at least in part on the one or more identified data annotation properties; output the modified data annotation pairs to at least one user; and generate a final collection of data annotation pairs by processing at least a portion of the modified data annotation pairs and user feedback received in response to the outputting of the modified data annotation pairs. 13 . The computer program product of claim 12 , wherein determining the one or more non-diffs comprises processing the at least a portion of the data annotation pairs for presence of similar length non-diffs in respective input data annotations and corresponding output data annotations. 14 . The computer program product of claim 12 , wherein determining the one or more non-diffs comprises processing the at least a portion of the data annotation pairs for presence of identical position placement of at least one non-diff in respective input data annotations and corresponding output data annotations. 15 . The computer program product of claim 12 , wherein determining the one or more non-diffs comprises processing the at least a portion of the data annotation pairs for presence of at least one non-diff of a same token type in respective input data annotations and corresponding output data annotations. 16 . The computer program product of claim 12 , wherein determining the one or more non-diffs comprises processing the at least a portion of the data annotation pairs for presence of one or more repeating characters within at least one non-diff in respective input data annotations and corresponding output data annotations. 17 . The computer program product of claim 12 , wherein the user feedback comprises at least one of acceptance of at least a portion of the one or more identified data annotation properties and rejection of at least a portion of the one or more identified data annotation properties. 18 . The computer program product of claim 12 , wherein the program instructions executable by a comput

Assignees

Inventors

Classifications

G06F16/2365
Ensuring data consistency and integrity · CPC title
G06F16/215Primary
Improving data quality; Data cleansing, e.g. de-duplication, removing invalid entries or correcting typographical errors · CPC title

Patent family

Related publications grouped by family.

View patent family 85774725

External sources

Frequently asked questions

Answers are generated from the same data shown on this page.

What does patent US12487976B2 cover?: Methods, systems, and computer program products for automatically improving data annotations by processing annotation properties and user feedback are provided herein. A computer-implemented method includes obtaining data annotation pairs, each comprising an input data annotation in a first format and a corresponding output data annotation in a second format; determining, within at least a port…
Who is the assignee on this patent?: IBM
What technology area does this patent fall under?: Primary CPC classification G06F16/215. Mapped technology areas include Physics.
When was this patent published?: Publication date Tue Dec 02 2025 00:00:00 GMT+0000 (Coordinated Universal Time) (B2). Legal status and post-grant events are not shown on this page.
What related patents are in patentsdb?: We list 3 related publications on this page (citations in our corpus or others sharing the same primary CPC).

How to read this patent

Abstract

First claim

Assignees

Inventors

Classifications

Patent family

External sources

Related patents

Automated quality assurance checks for improving the construction of natural language understanding systems

Method and apparatus for selecting among competing models in a tool for building natural language understanding models

Automated correction of natural language processing systems

Frequently asked questions