Generate Audio
Generate speech, music, or a sound effect from text (metered).
Pick mode first: text_to_speech (a voice saying your exact words —
prompt IS the script), text_to_music, or text_to_sfx. The model card on
GET /v1/media/models says which modes each model serves.
Managed voice models (elevenlabs-native-*, listed by
GET /v1/voices/capabilities) are a durable task lane instead: the answer is
an audio_generation task envelope (or, with quote: true, an
audio_generation_quote that starts nothing). wait: false returns it
queued; poll GET /v1/media/generate-audio/{task_id} no sooner than its
retry_after_ms, and stop it with POST /v1/media/generate-audio/{task_id}/cancel.
voice is a vox_ id, speech_to_speech converts a source_audio upload,
and an idempotency_key makes the call safe to repeat: the same key and
input return the same task (replayed: true), whatever its state. Everything
below this paragraph describes the other models.
Synchronous: the render runs inside this call and there is no wait: false
task lane. Speech returns in seconds. Music and sound effects can be asked
for at up to 600 seconds of output and may take correspondingly longer; a
render that outruns this call’s wait is asked to cancel and reported as a
retryable 503. Repeat that call with the same idempotency_key: it adopts
the existing provider job, so it returns the audio if the render finished
anyway and reports it cancelled if it did not — and it can never pay for a
second render. Only once it reports cancelled is a new request worth making
— and that one needs a NEW idempotency_key as well as a shorter duration or
fewer samples, since a changed payload under the old key is a 409
idempotency_conflict.
results carries one entry per delivered take (num_samples), applied
reports the knobs that actually ran, and adjustments names every field
that differed from the ask, with a reason — read them before describing the
output to a user.
Authorizations
API key from Settings > Developer > REST API
Headers
Calendar-dated API version pin. New integrations should pin 2026-05-01 to opt into the newest response shapes. For back-compat the server also accepts requests with no header and resolves them to the current default (today: 2026-04-12); that default advances on each sunset date. Any unsupported value returns 400 unsupported_version.
2026-04-12, 2026-05-01 "2026-05-01"
Body
Concrete audio model id; see GET /v1/media/models. Managed voice models (elevenlabs-native-*) are listed by GET /v1/voices/capabilities instead.
text_to_speech, text_to_music, or text_to_sfx. Stated rather than derived — every mode takes exactly text, so there is no input role to derive it from. The model card on GET /v1/media/models says which modes each model serves. Managed voice models also take speech_to_speech (aliases sts, voice_conversion) and require mode to equal the model's operation.
Sent to the selected model VERBATIM. For text_to_speech this IS the script — the words that get spoken, with no stage directions. For the other modes it describes the result: genre, instrumentation, mood, tempo, texture. Required for every mode except speech_to_speech.
Speech only. Normally a preset name from the mode's voices on GET /v1/media/models; on a mode whose card sets voice_is_free_form, voices is empty and this takes any provider voice name or cloned-voice id instead, checked for shape here and resolved by the provider. Omit for the mode's default. Managed voice models require a vox_ id from GET /v1/voices.
Music and sound effects only, and it SNAPS into the mode's envelope with the adjustment reported. Speech has no duration control — a value passed there is dropped and reported rather than rejected.
Distinct takes to render, where the mode has a sample axis (max_samples on GET /v1/media/models). EVERY sample is billed, and a mode with a billing floor bills each short take at the floor.
Model settings. For a managed voice model, exactly the control names its capabilities list; an unknown or out-of-range setting is a 422 unsupported_model_setting, never silently dropped.
Keys the provider-job checkpoint: a retried call with the same key resumes the existing render instead of paying for a duplicate. On a managed voice model it keys the durable execution: the same key and input return the same task (replayed: true), a different input is 409 idempotency_conflict, and keys never expire.
200Managed voice models only. true (the default) runs the generation inside this call for up to about 210 seconds and returns the task envelope, terminal when it finished in time; false returns the queued task envelope at once — poll its poll_url no sooner than retry_after_ms.
Managed voice models only. true validates the request and returns an audio_generation_quote (committed: false) with the usage estimate and a credit check. Nothing is started, held or charged, and the idempotency_key is ignored.
speech_to_speech only: the recording to convert, as a file_ id from POST /v1/uploads (MP3, WAV or AAC).
speech_to_speech only. on_failure deletes a source recording this user uploaded for this generation if the generation fails or is canceled; never (the default) keeps it.
on_failure, never Managed voice models with context stitching only: text or earlier tasks around this passage, so a regenerated passage matches its neighbours.
Managed voice models whose pronunciation mode is dictionary: pinned dictionary revisions, in precedence order.
speech_to_speech quotes without source_audio: the recording's length in seconds (above 0, at most 300).
Managed voice models only: the credit estimate the caller approved. A generation whose admission estimate is higher is refused (422 validation_failed, estimate_exceeds_approval) before anything is held or run.
Response
Successful Response
The response is of type Response Mediagenerateaudio · object.