Know what is in your training data
A measured report on any text or code dataset: duplicates, language integrity, personal data, licence compliance, syntax validity and overlap with evaluation sets. Before and after curation, in one document.
Text and code get their own checks
Code is not judged like prose. A broken source file, a leaked key or a copyleft header matters in code. A wrong language label or a damaged scan matters in text.
Text datasets
- Inventory: documents, tokens, length distribution, language and domain mix
- Exact and near duplicates
- Language identification against the supplied label
- Licence metadata against your licence policy
- Structural quality: repetition, encoding damage, fragments
- Detected personal data: e-mail, phone, IBAN, payment card, IP address
- Overlap with evaluation sets
Code datasets
- Inventory by programming language: files, tokens, lines
- Exact and near duplicate files
- Syntax validity: share of files that parse without errors
- Licence headers, including copyleft share
- Leaked credentials: keys, tokens, hard-coded passwords
- Generated, minified and vendored files
- Overlap with code evaluation sets
One figure, with the working shown
Each check turns a measured defect rate into a score from 0 to 100 against a stated tolerance. The index is the weighted mean, reported in bands from A to E, for the data before and after curation.
Tolerances and weights are printed in every report. The index summarises the measurements. It is not a guarantee of how a model trained on the data will perform.

A page from a report on synthetic test data with planted defects. The figures are not from a client dataset.
Five rules the report keeps
Numbers only
The report holds counts, rates and distributions. It contains no text from your data, and small groups are merged so no figure points at a handful of documents.
Local processing
Measurement runs on our machine or inside your environment. No dataset content is sent to an outside service.
Reproducible
Tool version, settings, tokenizer and a fingerprint of the data are recorded. The same input gives the same figures.
Detected, not absent
Scans report what was detected. A zero is not presented as proof that nothing is there, and every section states its known limits.
Your private evaluation sets
Your own benchmark can be checked for leakage. It is converted to a hashed index that cannot be turned back into text.
Four moments to measure
Before you buy data
Check a vendor's sample against what the offer claims.
Before you train
Find duplicates, personal data and benchmark leakage while they are still cheap to fix.
After curation
Show in figures what cleaning changed, for your own team or for an auditor.
When you sell data
Send buyers an independent quality report next to your dataset card.
What you receive
Send or host
You send the dataset under NDA, or we measure it inside your environment.
We measure
CurateLM runs the measurement. The tool itself is not handed over.
You get the report
A PDF report, an HTML version and a JSON file with every figure.
Questions about the benchmark
Does the report contain any of our data?
No. It contains aggregate numbers only. Detected e-mail addresses, keys and similar values are counted and never stored.
Can you check our data against our own private benchmark?
Yes. Your evaluation set is converted into a hashed index on the machine that runs the measurement. The index holds no benchmark text.
Is a good index a guarantee of model quality?
No. The report states what was measured and how. It is not legal advice, not a licence opinion, and not a prediction of model performance.
Which programming languages can be checked for syntax?
More than twenty, including Python, C, C++, Java, JavaScript, TypeScript, Rust, Go, SQL, COBOL and Fortran. Languages without a parser are reported as not checked, not as failed.
Can we get the benchmark without buying data or curation?
Yes. The benchmark is a separate service. It is also included with data curation projects as the before-and-after report.
Have a dataset measured
Tell us whether it is text, code or both, and roughly how large it is.
Request a benchmark