The record of public AI data
Know where your
data came from.
Every dataset your model learns from has a history. Archivum indexes what’s already public and keeps one consistent record of each dataset — origin, licensing, lineage, and exactly how much of it the source documents.
50 datasets indexed · 2 platforms · independent of every one of them
allenai · GitHub
natural-instructions
% documented
Lineage
The problem
You wouldn’t buy a used car without the history report.
You can see the dataset. You can’t see whether it was scraped legally, who touched it, when it was last real, or whether using it commercially will get you sued. Teams spend weeks hunting for data, then still can’t answer those questions.
Archivum is the history report.
Wrong answers ship
Outdated or dirty data reaches production, and the model hallucinates with confidence.
Licenses surface too late
Commercial restrictions get discovered after training, not before.
Weeks disappear
Searching, validating, and cleaning eats the time you meant to spend building.
How it works
From search to verified in an afternoon.
01
Search
One index across Hugging Face, Kaggle, GitHub, and academic sources.
Filter by domain, language, modality, licence, and minimum Documentation Coverage. Stop hunting across a dozen sites.
- 80%
natural-instructions
allenai
- 74%
Multi-Genre Natural Language Inference
nyu-mll
- 74%
Beans
AI-Lab-Makerere
- 73%
datasets
plotly
02
Verify
Open the passport. Read the report card.
Origin, contributors, update history, licensing, and the full lineage trail behind every transformation — with every claim labeled by how it was established.
allenai · GitHub
natural-instructions
- Source transparency94 · w35
- Community verification88 · w25
- Update frequency90 · w20
- Documentation quality91 · w20
03
Integrate
Pull it into your pipeline, provenance attached.
Download directly or export to LlamaIndex, LangChain, or your vector store. The provenance record travels with the data.
Documentation Coverage
A number anyone can recompute.
Documentation Coverage measures one thing: how much of a dataset’s provenance is documented at the source. Twenty-eight factual checks across four sections — each answering “was this present in the record?”, never “is this dataset good?”
coverage rules v1.0 · example record
- Origin & Sourcing7 checks · 94%
Where did this data come from, and who assembled it?
- Licensing & Terms7 checks · 88%
What terms did the publisher attach to reuse?
- Composition & Structure7 checks · 90%
What is actually inside, and how is it organised?
- Maintenance & Usage7 checks · 91%
Is it still maintained, and how is it being used?
Archivum retrieved the artifact itself from the platform API — a licence field, a file manifest, a commit history.
The publisher stated it in prose that Archivum retrieved but did not independently confirm.
Absent from the published metadata when Archivum checked. A fact about the record, not a defect in the data.
Every figure shows its arithmetic. Read how coverage is measured →
Coverage reflects what was present at the source at the time of the check. It describes documentation, not data quality, and is not legal advice.
The index
50 datasets. Every one graded.
Integrations
Archivum sits under the stack you already use.
Index datasets from the platforms where they already live. Export them, with provenance attached, into the tools you already build with.
Sources indexed
- Hugging Face
- Kaggle
- GitHub
- Zenodo
- Papers with Code
- Institutional repositories
Archivum · index, record, trace
Export targets
- LlamaIndex
- LangChain
- Pinecone
- Qdrant
- Weaviate
- Chroma
- S3
- Snowflake
Platform names are shown to indicate where Archivum indexes from and exports to. They do not indicate partnership, sponsorship, or endorsement. All trademarks belong to their respective owners.
Know what your model learned from.
Archivum is in early access. Join the waitlist and help shape what gets indexed first.
Publishing a dataset? Submit it →