Your unstructured data, made ready for AI
We run our curation pipeline on your documents, scans, audio and video: text extraction, personal-data scrubbing, deduplication, language detection and quality scoring. Processed in Germany and the EU under a data processing agreement, or on your own premises.
What goes in
- PDFs, including scanned pages
- Office documents, HTML and e-mail exports
- CSV files, logs and database exports
- Audio and video recordings
What comes out
- Structured JSONL or Parquet in an agreed schema
- Personal data removed or masked according to your policy
- Language, quality score and content hash on every record
- A data card and a before-and-after quality report
Eight steps from raw files to training data
The same pipeline that builds our own corpora. Each step can be switched on or off for your project.
Extract
Text is pulled out of documents. Scanned pages go through optical character recognition.
Transcribe
Audio and video are converted and transcribed to text.
Clean
Encoding damage, markup, page furniture and boilerplate are removed. Text is normalised.
Scrub personal data
Names, e-mail addresses, phone numbers, bank details and identifiers are detected and removed or masked.
Deduplicate
Exact copies and near copies are found and removed, so nothing is over-weighted in training.
Detect language
Four independent detectors vote. A document is accepted when at least two agree.
Score quality
Every record gets a quality score. Low-scoring records are flagged or removed, as you decide.
Structure and report
Output in your schema as JSONL or Parquet, with a data card and a benchmark report.
What we do not promise: automated personal-data detection does not find everything. We report detected rates and known limits, and for sensitive data we agree a review step with you before delivery.
Built for data that has to stay under control
Regulated organisations cannot hand data to a cloud service in a third country. This service is designed around that constraint.
Data processing agreement
An agreement under Art. 28 GDPR is signed before any data is touched. An NDA is standard.
Processing in Germany and the EU
Your data is processed on our own machines in Germany and the EU, or on yours. No transfer to third countries.
Encryption
AES-256-GCM encryption at rest. Keys are held per session and never written to disk.
Deletion
After delivery your data is deleted and you receive written confirmation.
On our machines or on yours
| On our machines | On your premises | |
|---|---|---|
| How it works | You send the data over an encrypted transfer. We process it in Germany and the EU and delete it after delivery. | We run the pipeline inside your environment, remotely on a machine you provide or on site. |
| Suited to | AI teams, software companies and data that may leave your network under contract. | Banks, insurers, healthcare, pharma and public bodies whose data may not leave the building. |
| Your data | Held encrypted for the duration of the project only. | Never leaves your infrastructure. |
Organisations with archives that AI cannot use yet
Banks and insurers
Contracts, correspondence and case files turned into training and retrieval data without personal details.
Healthcare and pharma
Reports, study documents and recorded speech prepared for domain models under strict data protection.
Public sector and legal
Scanned records, rulings and administrative documents made searchable and usable.
Consultancies and data vendors
The pipeline is available white-label, so you can offer curation to your own clients under your name.
Start small, then scale
Scoping
A call about your data, formats, volumes and data protection needs.
Pilot
We curate a sample of about 1,000 records and show the result with a quality report.
Quote
A fixed price based on what the pilot showed, not on a guess.
Full run
The whole dataset goes through the pipeline. You see progress and interim figures.
Delivery
Data, data card and report are delivered. Your source data is deleted and confirmed.
Questions about data curation
Does our data have to leave our premises?
No. The pipeline can run inside your environment, on a machine you provide. If you prefer, you send the data to us under a data processing agreement and we process it in Germany and the EU.
Do you sign a data processing agreement?
Yes. An agreement under Art. 28 GDPR is signed before we receive any data, together with an NDA.
Which file types can you process?
PDF including scans, Office documents, HTML, plain text, CSV and database exports, and audio and video files. Scans go through OCR and recordings through speech-to-text.
Which languages are supported?
Language detection covers a very wide range of languages, including low-resource ones, because the pipeline was built for a 150-language corpus. OCR and speech-to-text quality varies by language, which is one reason we start with a pilot.
What does it cost?
Pricing is per project and depends on volume, formats and the delivery model. The pilot gives both sides real numbers, and the quote follows from it.
Send us a description of your data
Formats, rough volume and what you want to use it for. We reply with a pilot proposal.
Request a pilot