The record of public AI data

Know where your
data came from.

Every dataset your model learns from has a history. Archivum indexes what’s already public and keeps one consistent record of each dataset — origin, licensing, lineage, and exactly how much of it the source documents.

50 datasets indexed · 2 platforms · independent of every one of them

Dataset recordExtensively documented

allenai · GitHub

natural-instructions

0

% documented

LicenseApache-2.0
Commercial termscommercial use permitted
Last updated2y ago
Records0

Lineage

The problem

You wouldn’t buy a used car without the history report.

You can see the dataset. You can’t see whether it was scraped legally, who touched it, when it was last real, or whether using it commercially will get you sued. Teams spend weeks hunting for data, then still can’t answer those questions.

Archivum is the history report.

  • Wrong answers ship

    Outdated or dirty data reaches production, and the model hallucinates with confidence.

  • Licenses surface too late

    Commercial restrictions get discovered after training, not before.

  • Weeks disappear

    Searching, validating, and cleaning eats the time you meant to spend building.

How it works

From search to verified in an afternoon.

01

Search

One index across Hugging Face, Kaggle, GitHub, and academic sources.

Filter by domain, language, modality, licence, and minimum Documentation Coverage. Stop hunting across a dozen sites.

medical text, commercial use
  • natural-instructions

    allenai

    80%
  • Multi-Genre Natural Language Inference

    nyu-mll

    74%
  • Beans

    AI-Lab-Makerere

    74%
  • datasets

    plotly

    73%

02

Verify

Open the passport. Read the report card.

Origin, contributors, update history, licensing, and the full lineage trail behind every transformation — with every claim labeled by how it was established.

allenai · GitHub

natural-instructions

80%
  • Source transparency94 · w35
  • Community verification88 · w25
  • Update frequency90 · w20
  • Documentation quality91 · w20
Documentedlineage 100% documented

03

Integrate

Pull it into your pipeline, provenance attached.

Download directly or export to LlamaIndex, LangChain, or your vector store. The provenance record travels with the data.

archivum · zsh
$archivum pull github-allenai-natural-instructions@55a3656
fingerprint verified
license Apache-2.0 · commercial allowed
provenance record attached
exported → llamaindex · github-allenai-natural-instructions.jsonl
$

Documentation Coverage

A number anyone can recompute.

Documentation Coverage measures one thing: how much of a dataset’s provenance is documented at the source. Twenty-eight factual checks across four sections — each answering “was this present in the record?”, never “is this dataset good?”

0% documented

coverage rules v1.0 · example record

  • Origin & Sourcing7 checks · 94%

    Where did this data come from, and who assembled it?

  • Licensing & Terms7 checks · 88%

    What terms did the publisher attach to reuse?

  • Composition & Structure7 checks · 90%

    What is actually inside, and how is it organised?

  • Maintenance & Usage7 checks · 91%

    Is it still maintained, and how is it being used?

Documented

Archivum retrieved the artifact itself from the platform API — a licence field, a file manifest, a commit history.

Reported

The publisher stated it in prose that Archivum retrieved but did not independently confirm.

Not found

Absent from the published metadata when Archivum checked. A fact about the record, not a defect in the data.

Every figure shows its arithmetic. Read how coverage is measured →

Coverage reflects what was present at the source at the time of the check. It describes documentation, not data quality, and is not legal advice.

The index

50 datasets. Every one graded.

allenai · GitHub

natural-instructions

80%

Expanding natural instructions

Apache-2.0commercial use permitted0 rows2y ago

nyu-mll · Hugging Face

Multi-Genre Natural Language Inference

74%

Dataset Card for Multi-Genre Natural Language Inference (MultiNLI) Dataset Summary The Multi-Genre Natural Language Inference (MultiNLI) corpus is a crowd-sourced collection of 433k sentence pairs annotated with textual entailment information. The corpus is modeled on the SNLI corpus, but differs in that covers a range of genres of spoken and written text, and supports a distinctive cross-genre generalization evaluation. The corpus served as the basis for the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-mll/multi_nli.

cc-by-3.0commercial use permitted412,349 rows2y ago

AI-Lab-Makerere · Hugging Face

Beans

74%

Dataset Card for Beans Dataset Summary Beans leaf dataset with images of diseased and health leaves. Supported Tasks and Leaderboards image-classification: Based on a leaf image, the goal of this task is to predict the disease type (Angular Leaf Spot and Bean Rust), if any. Languages English Dataset Structure Data Instances A sample from the training set is provided below: { 'image_file_path':… See the full description on the dataset page: https://huggingface.co/datasets/AI-Lab-Makerere/beans.

mitcommercial use permitted1,295 rows2y ago

plotly · GitHub

datasets

73%

Datasets used in Plotly examples and documentation

MITcommercial use permitted0 rows3mo ago

tatsu-lab · Hugging Face

Alpaca

72%

Dataset Card for Alpaca Dataset Summary Alpaca is a dataset of 52,000 instructions and demonstrations generated by OpenAI's text-davinci-003 engine. This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction better. The authors built on the data generation pipeline from Self-Instruct framework and made the following modifications: The text-davinci-003 engine to generate the instruction data… See the full description on the dataset page: https://huggingface.co/datasets/tatsu-lab/alpaca.

cc-by-nc-4.0non-commercial terms52,002 rows3y ago

nlphuji · Hugging Face

flickr30k

36%

Flickr30k Original paper: From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions Homepage: https://shannon.cs.illinois.edu/DenotationGraph/ Bibtex: @article{young2014image, title={From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions}, author={Young, Peter and Lai, Alice and Hodosh, Micah and Hockenmaier, Julia}, journal={Transactions of the… See the full description on the dataset page: https://huggingface.co/datasets/nlphuji/flickr30k.

Not statedterms not stated31,014 rows3y ago

Explore all 50 datasets

Integrations

Archivum sits under the stack you already use.

Index datasets from the platforms where they already live. Export them, with provenance attached, into the tools you already build with.

Sources indexed

  • Hugging Face
  • Kaggle
  • GitHub
  • Zenodo
  • Papers with Code
  • Institutional repositories

Archivum · index, record, trace

Export targets

  • LlamaIndex
  • LangChain
  • Pinecone
  • Qdrant
  • Weaviate
  • Chroma
  • S3
  • Snowflake

Platform names are shown to indicate where Archivum indexes from and exports to. They do not indicate partnership, sponsorship, or endorsement. All trademarks belong to their respective owners.

Know what your model learned from.

Archivum is in early access. Join the waitlist and help shape what gets indexed first.

Publishing a dataset? Submit it →