Pre-training data

Licensed corpora for model training, fine-tuning, and benchmarking. Books, scientific journals, and proprietary company data, cleared for AI use and delivered at the scale a training run needs.

Datasets by domain

What we license today, grouped by domain. Every dataset is cleared with its rights holder before delivery.

Science and engineering

Peer-reviewed science and research material across the whole STEM space.

Physics, chemistry, engineering
Research-grade content
Structured knowledge and explanations

Medical and healthcare

Clinical trial data and scientific datasets with a multitude of applications.

Clinical knowledge
Domain-specific language
High-precision datasets

Source code

A large body of code from private repos, suited to training coding agents.

Multiple languages
High test coverage
Private repos

Company data

Proprietary material from inside the business: decks, documents, and reports, licensed at the source and cleared for AI training.

Slides and presentations
Documents and reports
Manuals and specifications

Audio and voice data

Transcribed spoken audio, podcasts, and music to train speech and music models.

Music
Transcribed audio
Podcasts

World models and robotics

A wide set of data for training world models and robotics.

Human egocentric video
Multi-angle video
Physics for world models

Tell us what you need to train on

We license to order, including material that is not listed here. Tell us the domain, the modality, and the scale, and we will come back with what is available and on what terms.

Request sample access