Data quality benchmark

Know what is in your training data

A measured report on any text or code dataset: duplicates, language integrity, personal data, licence compliance, syntax validity and overlap with evaluation sets. Before and after curation, in one document.

What is measured

Text and code get their own checks

Code is not judged like prose. A broken source file, a leaked key or a copyleft header matters in code. A wrong language label or a damaged scan matters in text.

Text datasets

  • Inventory: documents, tokens, length distribution, language and domain mix
  • Exact and near duplicates
  • Language identification against the supplied label
  • Licence metadata against your licence policy
  • Structural quality: repetition, encoding damage, fragments
  • Detected personal data: e-mail, phone, IBAN, payment card, IP address
  • Overlap with evaluation sets

Code datasets

  • Inventory by programming language: files, tokens, lines
  • Exact and near duplicate files
  • Syntax validity: share of files that parse without errors
  • Licence headers, including copyleft share
  • Leaked credentials: keys, tokens, hard-coded passwords
  • Generated, minified and vendored files
  • Overlap with code evaluation sets
The quality index

One figure, with the working shown

Each check turns a measured defect rate into a score from 0 to 100 against a stated tolerance. The index is the weighted mean, reported in bands from A to E, for the data before and after curation.

AExcellent BGood CFair DWeak EPoor

Tolerances and weights are printed in every report. The index summarises the measurements. It is not a guarantee of how a model trained on the data will perform.

Summary page of a CurateLM data quality benchmark report showing the quality index, component scores and headline measurements

A page from a report on synthetic test data with planted defects. The figures are not from a client dataset.

How the report is made

Five rules the report keeps

Numbers only

The report holds counts, rates and distributions. It contains no text from your data, and small groups are merged so no figure points at a handful of documents.

Local processing

Measurement runs on our machine or inside your environment. No dataset content is sent to an outside service.

Reproducible

Tool version, settings, tokenizer and a fingerprint of the data are recorded. The same input gives the same figures.

Detected, not absent

Scans report what was detected. A zero is not presented as proof that nothing is there, and every section states its known limits.

Your private evaluation sets

Your own benchmark can be checked for leakage. It is converted to a hashed index that cannot be turned back into text.

When it helps

Four moments to measure

Before you buy data

Check a vendor's sample against what the offer claims.

Before you train

Find duplicates, personal data and benchmark leakage while they are still cheap to fix.

After curation

Show in figures what cleaning changed, for your own team or for an auditor.

When you sell data

Send buyers an independent quality report next to your dataset card.

Delivery

What you receive

Send or host

You send the dataset under NDA, or we measure it inside your environment.

We measure

CurateLM runs the measurement. The tool itself is not handed over.

You get the report

A PDF report, an HTML version and a JSON file with every figure.

Questions about the benchmark

Does the report contain any of our data?

No. It contains aggregate numbers only. Detected e-mail addresses, keys and similar values are counted and never stored.

Can you check our data against our own private benchmark?

Yes. Your evaluation set is converted into a hashed index on the machine that runs the measurement. The index holds no benchmark text.

Is a good index a guarantee of model quality?

No. The report states what was measured and how. It is not legal advice, not a licence opinion, and not a prediction of model performance.

Which programming languages can be checked for syntax?

More than twenty, including Python, C, C++, Java, JavaScript, TypeScript, Rust, Go, SQL, COBOL and Fortran. Languages without a parser are reported as not checked, not as failed.

Can we get the benchmark without buying data or curation?

Yes. The benchmark is a separate service. It is also included with data curation projects as the before-and-after report.

Have a dataset measured

Tell us whether it is text, code or both, and roughly how large it is.

Request a benchmark