API reference

Upload API

Push your own audio into Audioscrape — interviews, conference talks, internal recordings, customer calls — and have it transcribed, diarised, entity-extracted, and indexed alongside every podcast on the platform. One endpoint, two intake modes: send the file directly as multipart/form-data, or hand us a URL with application/json and we'll fetch it server-side.

Machine-readable spec: for the always-current OpenAPI definition, see the interactive Swagger UI or download /api/openapi.json.

Endpoint

The endpoint

All uploads target a single route. The Content-Type header decides the intake mode: multipart/form-data for raw bytes, application/json for URL-fetch.

POST /api/datasets/{id}/items

{id} is the numeric ID of an upload-type dataset you own (or are a member of via its parent workspace). Datasets created for podcast feeds will reject uploads — create a dataset with content_type: "uploads" first. The authenticated caller must be a workspace member; otherwise the API returns 404 (we deliberately don't disclose whether the dataset exists).

Intake mode 1

Multipart upload

Use multipart when the audio lives on the machine making the request — browser file picker, a CLI script streaming a local file, a desktop client. The file bytes travel in the audio form field.

POST /api/datasets/123/items · multipart/form-data
curl -X POST "https://www.audioscrape.com/api/datasets/123/items" \
     -H "Authorization: Bearer YOUR_API_KEY" \
     -F "audio=@/path/to/interview.mp3" \
     -F "title=Q3 customer interview — ACME Corp" \
     -F "description=30-min discovery call with the ACME product team" \
     -F "content_type=interview" \
     -F 'metadata={"num_speakers":2,"source_url":"https://example.com/call"}'
{
  "success": true,
  "item_id": 84217,
  "message": "Audio file queued for transcription with priority 100"
}

Form fields

Field Type Required Description
audio file yes The audio or video file. Accepted: mp3, wav, m4a, ogg, flac, aac, opus; video (mp4, mov, webm, mkv, avi) and phone/dictaphone containers (amr, 3gp, wma, aiff, caf, mpeg) are converted to audio at intake.
title text optional Display title. Defaults to the filename without extension if omitted.
description text optional Free-form description, surfaced in the item record and in search snippets.
content_type text optional What the audio is. Sets how many speakers we look for: phone_call and interview pin two, lecture, voicemail and monologue pin one, court_hearing uses the court layout. Omit for automatic detection (up to 20). The single most useful hint for clean speaker labels.
metadata JSON object optional Stored on the item and returned with it. Processing hints the pipeline reads: num_speakers (exact), or min_speakers and max_speakers (bounds); these beat the content_type default. Anything else is yours, e.g. source_url or case_name.
derive_title bool optional Default false on the API: the title you send, or the filename, is kept exactly. true lets the pipeline replace it with a title derived from the transcript.

Datasets configured with require_metadata = true will reject uploads that omit either title or description.

Intake mode 2

JSON URL upload

Use the JSON variant when the audio is already hosted somewhere on the public internet — an S3 presigned URL, a CDN object, your own static server. Audioscrape downloads it, stores a copy on its own object storage, and queues the same transcription pipeline. This mirrors what the transcribe_audio MCP tool does under the hood, so the two surfaces stay at parity.

POST /api/datasets/123/items · application/json
curl -X POST "https://www.audioscrape.com/api/datasets/123/items" \
     -H "Authorization: Bearer YOUR_API_KEY" \
     -H "Content-Type: application/json" \
     -d '{
       "audio_url": "https://files.example.com/recording.mp3",
       "title": "Founder fireside — March 2026",
       "description": "Recording from our internal all-hands"
     }'
{
  "success": true,
  "item_id": 84218,
  "message": "Audio fetched from URL and queued for transcription"
}

JSON body

Field Type Required Description
audio_url string yes Publicly reachable https:// URL. See SSRF rules below.
title string optional Display title. Defaults to the last path segment of the URL if omitted.
description string optional Free-form description stored on the item.
content_type string optional Same values and effect as the form field above.
metadata object optional Same as the form field above: num_speakers / min_speakers / max_speakers plus your own keys.
derive_title bool optional Default false: your title is kept exactly.

SSRF guards on audio_url

Because the server fetches whatever URL you provide, we apply strict filters to prevent the endpoint from being used as a reflector into internal infrastructure:

  • Scheme must be https — plain http and other schemes are rejected.
  • The hostname is resolved server-side; if any resolved IP is loopback, link-local, RFC 1918 private, broadcast, documentation, unspecified, or IPv6 unique-local / link-local, the request is rejected.
  • Fetch has a 120-second timeout and is capped at the dataset's effective file-size limit. Content-Length over the cap is rejected before the body is read.
  • HTTP status must be 2xx; non-success responses surface as an error to the caller.
  • The file extension (derived from the URL path) must match the dataset's allowed-formats list, same as multipart.

Heads up: the URL only needs to be reachable at the moment of upload. We store a copy on our own object storage immediately, so transcription, playback, and search keep working even if the source URL later expires or 404s.

Limits

File size and format limits

The effective size cap is the smaller of the dataset's max_file_size_mb setting and the account-plan cap. Free-tier accounts are hard-capped at 200 MB per upload regardless of dataset config; paid plans default to the dataset setting (500 MB out of the box, configurable higher). Files exceeding the cap return HTTP 413.

Accepted extensions: mp3, wav, m4a, ogg, flac, aac, opus stored as-is; mp4, mov, m4v, webm, mkv, avi (video) and amr, 3gp, wma, aiff, caf, mpeg, m4b (phone and dictaphone recorders) are converted to audio at intake. Olympus ds2/dss has no open decoder: export it as MP3 first. Datasets can widen this list in their settings.

Transcription minutes consumed by uploads count against your plan's monthly quota, billed on actual audio duration in one-second increments with a 30-second minimum per file. See /pricing for current plan limits.

Testing

Testing your integration

Use real spoken audio for upload tests. We host a small public-domain fixture for exactly this — a 13-second NASA recording of the Apollo 11 "one small step" moment, with two speakers:

curl -X POST "https://www.audioscrape.com/api/datasets/{id}/items" \
     -H "Authorization: Bearer YOUR_API_KEY" \
     -H "Content-Type: application/json" \
     -d '{"url": "https://www.audioscrape.com/static/fixtures/sample-speech-apollo11.mp3", "title": "Upload smoke test"}'

Don't test with silence or tones

Silent or tone-only files (sine waves, generated beeps, empty WAVs) are not errors: they complete normally with "no speech detected" and produce an empty transcript. If your smoke test asserts on transcript content, it will fail confusingly. Always use real speech.

Lifecycle

After upload

A successful upload returns immediately with an item_id. The actual processing happens asynchronously in the background:

  1. Storage. The audio is stored on durable object storage. For URL uploads the server fetches it first (SSRF-checked) and keeps its own copy, so processing and playback survive the source URL later expiring.
  2. Queued. The item is created with status queued and enters the processing queue. Datasets configured with auto_transcribe = false stop here until you trigger processing.
  3. Transcription & diarisation. Speech is transcribed and split by speaker on managed GPU infrastructure. Wall-clock time is roughly 1–3× real-time depending on queue depth.
  4. Enrichment & indexing. The item runs through the same downstream stages as every podcast episode — speaker labelling, named-entity extraction (people, organisations, places, products), and search indexing — then moves to published.

Poll the item to see status and grab the transcript once it's ready:

GET /api/items/{id}
curl "https://www.audioscrape.com/api/items/84217" \
     -H "Authorization: Bearer YOUR_API_KEY"

The item moves from queued through processing into the published state. Once published, the transcript and segments are available via the standard transcript and search endpoints, and the item appears in dataset listings under GET /api/datasets/{id}/items.

Response fields

Field Type Description
success boolean Always true on a 200 response.
item_id integer ID of the newly-created item. Use it with GET /api/items/{id}.
message string Human-readable status note (e.g. "queued for transcription").

Error responses

Status Meaning
400 Invalid request body or URL (bad JSON, non-https scheme, unparseable URL).
401 Missing or invalid API key.
403 Caller is not a member of the dataset's parent library, or transcription quota exhausted.
404 Dataset not found (also returned when caller lacks access — we don't disclose existence).
413 File exceeds the per-upload size cap on this plan.
415 Audio format not in the dataset's allowed-formats list.
422 Audio URL unreachable, returned non-2xx, or resolved to a forbidden internal address.