Arxiv Categories

arxiv-categories·7 columns·24 KB·CC0 1.0 for descriptive metadata; arXiv API Terms of Use for API/content constraints·updated 22d ago

Schema

category_idVARCHAR
category_nameVARCHAR
group_nameVARCHAR
archive_idVARCHAR
archive_nameVARCHAR
descriptionVARCHAR
source_urlVARCHAR

Get this dataset on your machine

subsets sync --bundle medium clones the mid-size tiers as queryable local tables — use --bundle large for the widest tier

pip install subsetsio
subsets sync --bundle small # or medium, or large
subsets serve # search + SQL at http://localhost:8080

Free and open — the data syncs straight from the public bucket, and search + SQL run entirely on your machine. See docs →

Query it locally

Full docs →
Bash
subsets serve   # local API on http://localhost:8080
curl -s -X POST "http://localhost:8080/query" \
  -H "Content-Type: application/json" \
  -d '{"sql": "SELECT * FROM \"arxiv-categories\" LIMIT 5"}'

No account, no API key — the data syncs straight from bundles.subsets.io and every query runs on your machine. See the local API reference.