Pre-training data
Licensed corpora for model training, fine-tuning, and benchmarking. Books, scientific journals, and proprietary company data, cleared for AI use and delivered at the scale a training run needs.
Datasets by domain
What we license today, grouped by domain. Every dataset is cleared with its rights holder before delivery.
Science and engineering
Peer-reviewed science and research material across the whole STEM space.
.jpg)
Medical and healthcare
Clinical trial data and scientific datasets with a multitude of applications.
.jpg)
Source code
A large body of code from private repos, suited to training coding agents.
.jpg)
Company data
Proprietary material from inside the business: decks, documents, and reports, licensed at the source and cleared for AI training.
.jpg)
Audio and voice data
Transcribed spoken audio, podcasts, and music to train speech and music models.

World models and robotics
A wide set of data for training world models and robotics.
.jpg)
Tell us what you need to train on
We license to order, including material that is not listed here. Tell us the domain, the modality, and the scale, and we will come back with what is available and on what terms.