Libraries, Datasets, and Items API
Libraries are published content catalogs — curated collections that have been shared for browsing and search, like the public Audioscrape corpus. The hierarchy is Library → Dataset → Item: a dataset groups related content inside a library; an item is one piece of audio — a podcast episode, an uploaded talk, a hearing recording, or any future format.
Looking for your own tenant data? Your workspace — including
every private dataset and upload — is served by the
Workspaces API
(GET /api/workspaces/{workspace}). Use the Libraries endpoints
below to discover and read published catalogs.
Items are universal. Each one has a source_type field
(podcast_episode, upload, youtube, ...)
so the same endpoints work across formats. The older
/api/podcasts and /api/episodes routes remain as
convenience aliases for the most common case, but new integrations should
prefer /api/items and /api/libraries — they
scale across content types.
Machine-readable spec: for the always-current OpenAPI definition, see the interactive Swagger UI or download /api/openapi.json.
List Libraries ¶
Returns every public library on Audioscrape, ordered by subscriber count. The
is_subscribed field on each entry tells you whether the calling
user is already a member.
# every public library, ordered by subscriber count
curl "https://www.audioscrape.com/api/libraries" \
-H "Authorization: Bearer YOUR_API_KEY"{
"libraries": [
{
"slug": "audioscrape",
"name": "Audioscrape Public Library",
"description": "The public Audioscrape corpus — podcasts, talks, and hearings.",
"curator": "Audioscrape",
"subscriber_count": 142,
"dataset_count": 7,
"episode_count": 1834,
"is_subscribed": true
}
],
"total": 1
}Response fields
Get Library ¶
Returns one library plus the datasets the calling user can see. Public
datasets are always returned; workspace-scoped datasets are returned only
to members. Private libraries return 404 to non-members.
# one library + the datasets you can see
curl "https://www.audioscrape.com/api/libraries/audioscrape" \
-H "Authorization: Bearer YOUR_API_KEY"{
"slug": "audioscrape",
"name": "Audioscrape Public Library",
"description": "The public Audioscrape corpus — podcasts, talks, and hearings.",
"curator": "Audioscrape",
"subscriber_count": 142,
"dataset_count": 7,
"episode_count": 1834,
"is_subscribed": true,
"datasets": [
{
"id": 12,
"slug": "ai-research",
"name": "AI Research Podcasts",
"description": "Long-form interviews with AI researchers.",
"visibility": "public",
"podcast_count": 14,
"episode_count": 412
}
]
}Path parameters
Dataset fields
Get Dataset ¶
Returns a dataset plus a reference to its parent library, in a single response. Saves a round-trip when you have a dataset ID from a search result or item record and need the surrounding context.
# dataset + parent library in one call
curl "https://www.audioscrape.com/api/datasets/12" \
-H "Authorization: Bearer YOUR_API_KEY"{
"id": 12,
"slug": "ai-research",
"name": "AI Research Podcasts",
"description": "Long-form interviews with AI researchers.",
"visibility": "public",
"podcast_count": 14,
"episode_count": 412,
"library_slug": "audioscrape",
"library_name": "Audioscrape Public Library"
}Path parameters
List Items ¶
Universal listing across content types — podcast episodes, uploads, and
future formats. Filter by source_type, dataset_id,
or library (slug). Results are newest-first. Use this in place
of /api/episodes when your integration needs to work with
multiple content types.
# newest-first items, filtered to one library and source type
curl "https://www.audioscrape.com/api/items?library=audioscrape&source_type=podcast_episode&limit=20" \
-H "Authorization: Bearer YOUR_API_KEY"{
"items": [
{
"id": 98765,
"title": "Scaling Laws for Language Models",
"slug": "scaling-laws-for-language-models",
"source_type": "podcast_episode",
"publish_date": "2026-05-14",
"duration_seconds": 4280,
"has_transcript": true,
"podcast": {
"id": 321,
"title": "AI Research Today",
"slug": "ai-research-today"
},
"dataset_id": 12
}
],
"total": 1,
"limit": 20,
"offset": 0
}Query parameters
Item fields
Quota note: every Libraries / Datasets / Items request counts against your data-call quota. See the pricing page for per-plan limits.
Get Item ¶
Returns one item plus a reference to its parent podcast (when applicable).
Heavy fields are opt-in via ?include= — pass
transcript, entities, or both to load them in the
same response. This avoids paying the cost of a 500 KB+ transcript
when you only need metadata. Equivalent to
GET /api/episodes/{id} for podcast episodes, but works
uniformly across all source types.
# metadata plus opt-in heavy fields in one response
curl "https://www.audioscrape.com/api/items/98765?include=transcript,entities" \
-H "Authorization: Bearer YOUR_API_KEY"{
"episode": {
"id": 98765,
"title": "Scaling Laws for Language Models",
"slug": "scaling-laws-for-language-models",
"description": "A deep dive into compute-optimal training...",
"publish_date": "2026-05-14",
"duration_seconds": 4280,
"enclosure_url": "https://traffic.example.com/episode-98765.mp3"
},
"podcast": {
"id": 321,
"title": "AI Research Today",
"slug": "ai-research-today",
"image_url": "https://images.audioscrape.com/podcasts/321.jpg"
},
"transcript": {
"segments": [
{ "text": "Welcome back to the show.", "start": 0.0, "end": 2.4, "speaker": "SPEAKER_00" }
],
"speakers": ["SPEAKER_00", "SPEAKER_01"],
"total_segments": 1842
},
"entities": [
{ "name": "OpenAI", "entity_type": "organization", "slug": "openai", "mention_count": 12 }
]
}Path & query parameters
Get Item Transcript ¶
Returns the transcript on its own — segments plus speaker list — for any item, regardless of source type. Use this when you already have item metadata and only need the transcript payload.
# transcript only — segments + speakers
curl "https://www.audioscrape.com/api/items/98765/transcript" \
-H "Authorization: Bearer YOUR_API_KEY"{
"segments": [
{ "text": "Welcome back to the show.", "start": 0.0, "end": 2.4, "speaker": "SPEAKER_00" },
{ "text": "Today we’re talking about scaling laws.", "start": 2.5, "end": 5.1, "speaker": "SPEAKER_00" }
],
"speakers": ["SPEAKER_00", "SPEAKER_01"],
"total_segments": 1842
}