Licensed training data for LLMs
Text, code and biomedical corpora collected under one strict licence policy, scrubbed for personal data, deduplicated and scored. Delivered as JSONL or Parquet with a dataset card.
Over 5.5 billion tokens in 150+ languages
A text corpus built for the languages and subjects that general web crawls cover poorly. The focus is on rare and low-resource languages and on institutional and specialist material: science, law, government, medicine, finance and engineering.
- Tokens
- 5.5 billion+, growing daily
- Languages
- 150+, including many low-resource languages
- Domain niches
- 140
- Licences
- Public domain, CC0, CC BY
- Formats
- JSONL, Apache Parquet
Licence-verified code, including the legacy languages
Banks, insurers and public bodies still run on COBOL, Fortran, JCL and PL/I, and models that modernise those systems need training data in those languages. This corpus combines legacy enterprise code with modern languages. Only permissive licences are accepted, and the licence is checked for every repository.
- Tokens
- 2 billion+, growing daily
- Legacy languages
- COBOL, Fortran, JCL, PL/I
- Modern languages
- Rust, Go, Java, C, C++, JavaScript, TypeScript, MATLAB, Perl, Shell, SQL
- Licences
- MIT, Apache-2.0, BSD, ISC, Unlicense. No GPL or AGPL
- Formats
- JSONL, Apache Parquet
Over 700 million tokens of biomedical text, plus structured records
Biomedical text across 31 niches and 33 languages, from pharmacology and neurology to tropical medicine and biochemistry. Structured records are available alongside it for drug discovery and clinical work. Nothing in this corpus is patient-level data.
Not offered: patient-level data for digital twins. Adverse-event reports are de-identified regulatory filings, not a patient cohort, and must not be analysed as one.
- Tokens
- 700 million+, growing
- Documents
- 2,800+
- Domain niches
- 31, including medical, pharmacology, neurology, tropical medicine, biochemistry, immunology, epidemiology
- Languages
- 33
- Structured records
- Regulatory drug labelling, de-identified adverse-event reports, clinical trial registry records, measured binding affinities, gene and pathway annotations
- Use cases
- Target discovery, drug mechanisms, molecule optimisation, biomarker prediction, patient selection
Languages and subjects in the text corpus
Training data for languages and subjects that are hard to find elsewhere. Volumes differ a lot by language, so ask for the dataset card to see what is available for the ones you need.
Languages include
Abkhazian · Acehnese · Afrikaans · Albanian · Amharic · Ancient Greek · Arabic · Armenian · Avar · Aymara · Azerbaijani · Bambara · Bashkir · Basque · Belarusian · Bengali · Bislama · Breton · Bulgarian · Burmese · Catalan · Cebuano · Chamorro · Chechen · Chichewa · Chinese · Cornish · Croatian · Czech · Danish · Dutch · Dzongkha · English · Esperanto · Estonian · Fijian · Finnish · French · Fula · Galician · Ge'ez · Georgian · German · Gothic · Greek · Guaraní · Gujarati · Hausa · Hawaiian · Hebrew · Hindi · Hungarian · Icelandic · Igbo · Indonesian · Inuktitut · Irish · Italian · Japanese · Javanese · Kazakh · Khmer · Kinyarwanda · Korean · Kyrgyz · Lao · Latin · Latvian · Lezgian · Lingala · Lithuanian · Lower Sorbian · Luganda · Macedonian · Maithili · Malay · Malayalam · Maltese · Manx · Māori · Marathi · Marshallese · Mon · Mongolian · Nahuatl · Nepali · Norwegian · Occitan · Odia · Oromo · Ossetian · Persian · Polish · Portuguese · Punjabi · Quechua · Romanian · Russian · Samoan · Sanskrit · Serbian · Sesotho · Shona · Sindhi · Sinhala · Slovak · Slovenian · Somali · Spanish · Sundanese · Swahili · Swedish · Tagalog · Tajik · Tamil · Tatar · Telugu · Thai · Tibetan · Tigrinya · Tok Pisin · Tongan · Tsonga · Tswana · Turkish · Turkmen · Ukrainian · Urdu · Uyghur · Uzbek · Vietnamese · Welsh · Wolof · Xhosa · Yoruba · Zulu
Subjects include
Science and research · Law, courts and legislation · Government and parliamentary records · Medicine · Traditional medicine, Ayurveda and Islamic medicine (Tibb Nabawi) · Indigenous knowledge · Banking, finance and tax · Engineering and aeronautics · Agriculture · Architecture · Astronomy and chemistry · Linguistics · Theology and humanities · Education · Legacy IT systems
What every record carries
Each record is delivered with its licence, language, domain, quality score, token count and a content hash, so you can filter, audit and reproduce your selection.
One rule for everything we collect
A record is admitted only if its own licence allows commercial use and redistribution. A record with no licence is refused, not assumed to be open.
Accepted
- Public domain and CC0
- CC BY, with the attribution delivered alongside the data
- Code: MIT, Apache-2.0, BSD, ISC, Unlicense
Refused
- Non-commercial (NC) and no-derivatives (ND) licences
- Share-alike (SA) licences
- GPL, AGPL and other copyleft code licences
- Anything with a missing or unclear licence
From first look to delivery
Dataset card
You receive the dataset card with volumes, languages, niches and licence mix.
Sample
A sample in the delivery format shows the schema and the quality.
NDA and agreement
We sign an NDA, then agree scope, licence terms and price.
Delivery
Encrypted transfer of JSONL or Parquet files with documentation.
Questions buyers ask
Can we see a sample before we sign anything?
Yes. The dataset card is available on request, and a short sample in the delivery format follows. Larger evaluation packs are shared under NDA.
Do you disclose where the data comes from?
No. Sourcing is confidential. Each record carries its licence, and attribution is delivered wherever a licence requires it. CurateLM warrants licence compliance in the agreement.
Is personal data removed?
Text is scrubbed with automated detection for names, e-mail addresses, phone numbers and similar identifiers. Automated detection does not find everything, so we report detected rates and do not claim that none remains.
Can we license a single language or niche?
Yes. Every record is labelled with language and niche, so a slice can be cut by language, subject, licence class or quality score.
Do you provide a quality report with the data?
Yes. A benchmark report can be delivered with the dataset. It measures duplicates, detected personal data, licence mix and overlap with evaluation sets.
Ask for the dataset card
Tell us which languages, subjects or programming languages you need.
Request the dataset card