B2B data curation

Your unstructured data, made ready for AI

We run our curation pipeline on your documents, scans, audio and video: text extraction, personal-data scrubbing, deduplication, language detection and quality scoring. Processed in Germany and the EU under a data processing agreement, or on your own premises.

What goes in

  • PDFs, including scanned pages
  • Office documents, HTML and e-mail exports
  • CSV files, logs and database exports
  • Audio and video recordings

What comes out

  • Structured JSONL or Parquet in an agreed schema
  • Personal data removed or masked according to your policy
  • Language, quality score and content hash on every record
  • A data card and a before-and-after quality report
The pipeline

Eight steps from raw files to training data

The same pipeline that builds our own corpora. Each step can be switched on or off for your project.

Extract

Text is pulled out of documents. Scanned pages go through optical character recognition.

Tesseract OCR

Transcribe

Audio and video are converted and transcribed to text.

ffmpegWhisper

Clean

Encoding damage, markup, page furniture and boilerplate are removed. Text is normalised.

pandas

Scrub personal data

Names, e-mail addresses, phone numbers, bank details and identifiers are detected and removed or masked.

NERpattern rules

Deduplicate

Exact copies and near copies are found and removed, so nothing is over-weighted in training.

content hashMinHash

Detect language

Four independent detectors vote. A document is accepted when at least two agree.

GlotLIDfastTextLingualangdetect

Score quality

Every record gets a quality score. Low-scoring records are flagged or removed, as you decide.

Structure and report

Output in your schema as JSONL or Parquet, with a data card and a benchmark report.

JSONLParquet

What we do not promise: automated personal-data detection does not find everything. We report detected rates and known limits, and for sensitive data we agree a review step with you before delivery.

Data protection

Built for data that has to stay under control

Regulated organisations cannot hand data to a cloud service in a third country. This service is designed around that constraint.

Data processing agreement

An agreement under Art. 28 GDPR is signed before any data is touched. An NDA is standard.

Processing in Germany and the EU

Your data is processed on our own machines in Germany and the EU, or on yours. No transfer to third countries.

Encryption

AES-256-GCM encryption at rest. Keys are held per session and never written to disk.

Deletion

After delivery your data is deleted and you receive written confirmation.

Two ways to run it

On our machines or on yours

On our machinesOn your premises
How it worksYou send the data over an encrypted transfer. We process it in Germany and the EU and delete it after delivery.We run the pipeline inside your environment, remotely on a machine you provide or on site.
Suited toAI teams, software companies and data that may leave your network under contract.Banks, insurers, healthcare, pharma and public bodies whose data may not leave the building.
Your dataHeld encrypted for the duration of the project only.Never leaves your infrastructure.
Who it is for

Organisations with archives that AI cannot use yet

Banks and insurers

Contracts, correspondence and case files turned into training and retrieval data without personal details.

Healthcare and pharma

Reports, study documents and recorded speech prepared for domain models under strict data protection.

Public sector and legal

Scanned records, rulings and administrative documents made searchable and usable.

Consultancies and data vendors

The pipeline is available white-label, so you can offer curation to your own clients under your name.

How an engagement runs

Start small, then scale

Scoping

A call about your data, formats, volumes and data protection needs.

Pilot

We curate a sample of about 1,000 records and show the result with a quality report.

Quote

A fixed price based on what the pilot showed, not on a guess.

Full run

The whole dataset goes through the pipeline. You see progress and interim figures.

Delivery

Data, data card and report are delivered. Your source data is deleted and confirmed.

Questions about data curation

Does our data have to leave our premises?

No. The pipeline can run inside your environment, on a machine you provide. If you prefer, you send the data to us under a data processing agreement and we process it in Germany and the EU.

Do you sign a data processing agreement?

Yes. An agreement under Art. 28 GDPR is signed before we receive any data, together with an NDA.

Which file types can you process?

PDF including scans, Office documents, HTML, plain text, CSV and database exports, and audio and video files. Scans go through OCR and recordings through speech-to-text.

Which languages are supported?

Language detection covers a very wide range of languages, including low-resource ones, because the pipeline was built for a 150-language corpus. OCR and speech-to-text quality varies by language, which is one reason we start with a pilot.

What does it cost?

Pricing is per project and depends on volume, formats and the delivery model. The pilot gives both sides real numbers, and the quote follows from it.

Send us a description of your data

Formats, rough volume and what you want to use it for. We reply with a pilot proposal.

Request a pilot