Datasets

Licensed training data for LLMs

Text, code and biomedical corpora collected under one strict licence policy, scrubbed for personal data, deduplicated and scored. Delivered as JSONL or Parquet with a dataset card.

01 · Multilingual text

Over 5.5 billion tokens in 150+ languages

A text corpus built for the languages and subjects that general web crawls cover poorly. The focus is on rare and low-resource languages and on institutional and specialist material: science, law, government, medicine, finance and engineering.

Tokens
5.5 billion+, growing daily
Languages
150+, including many low-resource languages
Domain niches
140
Licences
Public domain, CC0, CC BY
Formats
JSONL, Apache Parquet
02 · Source code

Licence-verified code, including the legacy languages

Banks, insurers and public bodies still run on COBOL, Fortran, JCL and PL/I, and models that modernise those systems need training data in those languages. This corpus combines legacy enterprise code with modern languages. Only permissive licences are accepted, and the licence is checked for every repository.

Tokens
2 billion+, growing daily
Legacy languages
COBOL, Fortran, JCL, PL/I
Modern languages
Rust, Go, Java, C, C++, JavaScript, TypeScript, MATLAB, Perl, Shell, SQL
Licences
MIT, Apache-2.0, BSD, ISC, Unlicense. No GPL or AGPL
Formats
JSONL, Apache Parquet
03 · Biomedical and life science

Over 700 million tokens of biomedical text, plus structured records

Biomedical text across 31 niches and 33 languages, from pharmacology and neurology to tropical medicine and biochemistry. Structured records are available alongside it for drug discovery and clinical work. Nothing in this corpus is patient-level data.

Not offered: patient-level data for digital twins. Adverse-event reports are de-identified regulatory filings, not a patient cohort, and must not be analysed as one.

Tokens
700 million+, growing
Documents
2,800+
Domain niches
31, including medical, pharmacology, neurology, tropical medicine, biochemistry, immunology, epidemiology
Languages
33
Structured records
Regulatory drug labelling, de-identified adverse-event reports, clinical trial registry records, measured binding affinities, gene and pathway annotations
Use cases
Target discovery, drug mechanisms, molecule optimisation, biomarker prediction, patient selection
Coverage

Languages and subjects in the text corpus

Training data for languages and subjects that are hard to find elsewhere. Volumes differ a lot by language, so ask for the dataset card to see what is available for the ones you need.

Languages include

Abkhazian · Acehnese · Afrikaans · Albanian · Amharic · Ancient Greek · Arabic · Armenian · Avar · Aymara · Azerbaijani · Bambara · Bashkir · Basque · Belarusian · Bengali · Bislama · Breton · Bulgarian · Burmese · Catalan · Cebuano · Chamorro · Chechen · Chichewa · Chinese · Cornish · Croatian · Czech · Danish · Dutch · Dzongkha · English · Esperanto · Estonian · Fijian · Finnish · French · Fula · Galician · Ge'ez · Georgian · German · Gothic · Greek · Guaraní · Gujarati · Hausa · Hawaiian · Hebrew · Hindi · Hungarian · Icelandic · Igbo · Indonesian · Inuktitut · Irish · Italian · Japanese · Javanese · Kazakh · Khmer · Kinyarwanda · Korean · Kyrgyz · Lao · Latin · Latvian · Lezgian · Lingala · Lithuanian · Lower Sorbian · Luganda · Macedonian · Maithili · Malay · Malayalam · Maltese · Manx · Māori · Marathi · Marshallese · Mon · Mongolian · Nahuatl · Nepali · Norwegian · Occitan · Odia · Oromo · Ossetian · Persian · Polish · Portuguese · Punjabi · Quechua · Romanian · Russian · Samoan · Sanskrit · Serbian · Sesotho · Shona · Sindhi · Sinhala · Slovak · Slovenian · Somali · Spanish · Sundanese · Swahili · Swedish · Tagalog · Tajik · Tamil · Tatar · Telugu · Thai · Tibetan · Tigrinya · Tok Pisin · Tongan · Tsonga · Tswana · Turkish · Turkmen · Ukrainian · Urdu · Uyghur · Uzbek · Vietnamese · Welsh · Wolof · Xhosa · Yoruba · Zulu

Subjects include

Science and research · Law, courts and legislation · Government and parliamentary records · Medicine · Traditional medicine, Ayurveda and Islamic medicine (Tibb Nabawi) · Indigenous knowledge · Banking, finance and tax · Engineering and aeronautics · Agriculture · Architecture · Astronomy and chemistry · Linguistics · Theology and humanities · Education · Legacy IT systems

What every record carries

Each record is delivered with its licence, language, domain, quality score, token count and a content hash, so you can filter, audit and reproduce your selection.

{ "text": "…", "language": "de", "niche": "medical", "license": "CC0-1.0", "quality_score": 0.91, "token_count": 1842, "text_hash": "9f2c…e41a" }
Licence policy

One rule for everything we collect

A record is admitted only if its own licence allows commercial use and redistribution. A record with no licence is refused, not assumed to be open.

Accepted

  • Public domain and CC0
  • CC BY, with the attribution delivered alongside the data
  • Code: MIT, Apache-2.0, BSD, ISC, Unlicense

Refused

  • Non-commercial (NC) and no-derivatives (ND) licences
  • Share-alike (SA) licences
  • GPL, AGPL and other copyleft code licences
  • Anything with a missing or unclear licence
How licensing works

From first look to delivery

Dataset card

You receive the dataset card with volumes, languages, niches and licence mix.

Sample

A sample in the delivery format shows the schema and the quality.

NDA and agreement

We sign an NDA, then agree scope, licence terms and price.

Delivery

Encrypted transfer of JSONL or Parquet files with documentation.

Questions buyers ask

Can we see a sample before we sign anything?

Yes. The dataset card is available on request, and a short sample in the delivery format follows. Larger evaluation packs are shared under NDA.

Do you disclose where the data comes from?

No. Sourcing is confidential. Each record carries its licence, and attribution is delivered wherever a licence requires it. CurateLM warrants licence compliance in the agreement.

Is personal data removed?

Text is scrubbed with automated detection for names, e-mail addresses, phone numbers and similar identifiers. Automated detection does not find everything, so we report detected rates and do not claim that none remains.

Can we license a single language or niche?

Yes. Every record is labelled with language and niche, so a slice can be cut by language, subject, licence class or quality score.

Do you provide a quality report with the data?

Yes. A benchmark report can be delivered with the dataset. It measures duplicates, detected personal data, licence mix and overlap with evaluation sets.

Ask for the dataset card

Tell us which languages, subjects or programming languages you need.

Request the dataset card