Upload API
Push your own audio into Audioscrape — interviews, conference talks, internal recordings, customer
calls — and have it transcribed, diarised, entity-extracted, and indexed alongside every podcast on
the platform. One endpoint, two intake modes: send the file directly as multipart/form-data,
or hand us a URL with application/json and we'll fetch it server-side.
Machine-readable spec: for the always-current OpenAPI definition, see the interactive Swagger UI or download /api/openapi.json.
The endpoint ¶
All uploads target a single route. The Content-Type header decides the intake mode:
multipart/form-data for raw bytes, application/json for URL-fetch.
{id} is the numeric ID of an upload-type dataset you own (or are a member of
via its parent workspace). Datasets created for podcast feeds will reject uploads — create a dataset
with content_type: "uploads" first. The authenticated caller must be a workspace member;
otherwise the API returns 404 (we deliberately don't disclose whether the dataset exists).
Multipart upload ¶
Use multipart when the audio lives on the machine making the request — browser file picker, a CLI
script streaming a local file, a desktop client. The file bytes travel in the audio form
field.
curl -X POST "https://www.audioscrape.com/api/datasets/123/items" \
-H "Authorization: Bearer YOUR_API_KEY" \
-F "audio=@/path/to/interview.mp3" \
-F "title=Q3 customer interview — ACME Corp" \
-F "description=30-min discovery call with the ACME product team" \
-F "content_type=interview" \
-F 'metadata={"num_speakers":2,"source_url":"https://example.com/call"}'{
"success": true,
"item_id": 84217,
"message": "Audio file queued for transcription with priority 100"
}Form fields
Datasets configured with require_metadata = true will reject uploads that omit either
title or description.
JSON URL upload ¶
Use the JSON variant when the audio is already hosted somewhere on the public internet — an S3
presigned URL, a CDN object, your own static server. Audioscrape downloads it, stores a copy on its own
object storage, and queues the same transcription pipeline. This mirrors what the
transcribe_audio MCP tool does under the hood, so the two surfaces stay at parity.
curl -X POST "https://www.audioscrape.com/api/datasets/123/items" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"audio_url": "https://files.example.com/recording.mp3",
"title": "Founder fireside — March 2026",
"description": "Recording from our internal all-hands"
}'{
"success": true,
"item_id": 84218,
"message": "Audio fetched from URL and queued for transcription"
}JSON body
SSRF guards on audio_url
Because the server fetches whatever URL you provide, we apply strict filters to prevent the endpoint from being used as a reflector into internal infrastructure:
- Scheme must be
https— plainhttpand other schemes are rejected. - The hostname is resolved server-side; if any resolved IP is loopback, link-local, RFC 1918 private, broadcast, documentation, unspecified, or IPv6 unique-local / link-local, the request is rejected.
- Fetch has a 120-second timeout and is capped at the dataset's effective file-size limit.
Content-Lengthover the cap is rejected before the body is read. - HTTP status must be 2xx; non-success responses surface as an error to the caller.
- The file extension (derived from the URL path) must match the dataset's allowed-formats list, same as multipart.
Heads up: the URL only needs to be reachable at the moment of upload. We store a copy on our own object storage immediately, so transcription, playback, and search keep working even if the source URL later expires or 404s.
File size and format limits ¶
The effective size cap is the smaller of the dataset's max_file_size_mb setting and the
account-plan cap. Free-tier accounts are hard-capped at 200 MB per upload regardless
of dataset config; paid plans default to the dataset setting (500 MB out of the box, configurable
higher). Files exceeding the cap return HTTP 413.
Accepted extensions: mp3, wav, m4a, ogg, flac,
aac, opus stored as-is; mp4, mov, m4v,
webm, mkv, avi (video) and amr, 3gp,
wma, aiff, caf, mpeg, m4b (phone and
dictaphone recorders) are converted to audio at intake. Olympus ds2/dss
has no open decoder: export it as MP3 first. Datasets can widen this list in their settings.
Transcription minutes consumed by uploads count against your plan's monthly quota, billed on actual audio duration in one-second increments with a 30-second minimum per file. See /pricing for current plan limits.
Testing your integration ¶
Use real spoken audio for upload tests. We host a small public-domain fixture for exactly this — a 13-second NASA recording of the Apollo 11 "one small step" moment, with two speakers:
curl -X POST "https://www.audioscrape.com/api/datasets/{id}/items" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://www.audioscrape.com/static/fixtures/sample-speech-apollo11.mp3", "title": "Upload smoke test"}'Don't test with silence or tones
Silent or tone-only files (sine waves, generated beeps, empty WAVs) are not errors: they complete normally with "no speech detected" and produce an empty transcript. If your smoke test asserts on transcript content, it will fail confusingly. Always use real speech.
After upload ¶
A successful upload returns immediately with an item_id. The actual processing happens
asynchronously in the background:
- Storage. The audio is stored on durable object storage. For URL uploads the server fetches it first (SSRF-checked) and keeps its own copy, so processing and playback survive the source URL later expiring.
-
Queued. The item is created with status
queuedand enters the processing queue. Datasets configured withauto_transcribe = falsestop here until you trigger processing. - Transcription & diarisation. Speech is transcribed and split by speaker on managed GPU infrastructure. Wall-clock time is roughly 1–3× real-time depending on queue depth.
-
Enrichment & indexing. The item runs through the same downstream stages as every
podcast episode — speaker labelling, named-entity extraction (people, organisations, places,
products), and search indexing — then moves to
published.
Poll the item to see status and grab the transcript once it's ready:
curl "https://www.audioscrape.com/api/items/84217" \
-H "Authorization: Bearer YOUR_API_KEY"
The item moves from queued through processing into the published state. Once
published, the transcript and segments are available via the standard transcript and search endpoints,
and the item appears in dataset listings under
GET /api/datasets/{id}/items.