Layovelle

Generate Audio

Documentation

POST
Generate Audio

Authorizations

Authorization
string
header
required

API key from Settings > Developer > REST API

Headers

Layovelle-Version
enum<string>
default:2026-05-01

Calendar-dated API version pin. New integrations should pin 2026-05-01 to opt into the newest response shapes. For back-compat the server also accepts requests with no header and resolves them to the current default (today: 2026-04-12); that default advances on each sunset date. Any unsupported value returns 400 unsupported_version.

Available options:
2026-04-12,
2026-05-01
Example:

"2026-05-01"

Body

application/json
model
string
required

Concrete audio model id; see GET /v1/media/models. Managed voice models (elevenlabs-native-*) are listed by GET /v1/voices/capabilities instead.

mode
string
required

text_to_speech, text_to_music, or text_to_sfx. Stated rather than derived — every mode takes exactly text, so there is no input role to derive it from. The model card on GET /v1/media/models says which modes each model serves. Managed voice models also take speech_to_speech (aliases sts, voice_conversion) and require mode to equal the model's operation.

prompt
string | null

Sent to the selected model VERBATIM. For text_to_speech this IS the script — the words that get spoken, with no stage directions. For the other modes it describes the result: genre, instrumentation, mood, tempo, texture. Required for every mode except speech_to_speech.

voice
string | null

Speech only. Normally a preset name from the mode's voices on GET /v1/media/models; on a mode whose card sets voice_is_free_form, voices is empty and this takes any provider voice name or cloned-voice id instead, checked for shape here and resolved by the provider. Omit for the mode's default. Managed voice models require a vox_ id from GET /v1/voices.

duration_seconds

Music and sound effects only, and it SNAPS into the mode's envelope with the adjustment reported. Speech has no duration control — a value passed there is dropped and reported rather than rejected.

num_samples
integer | null

Distinct takes to render, where the mode has a sample axis (max_samples on GET /v1/media/models). EVERY sample is billed, and a mode with a billing floor bills each short take at the floor.

model_params
Model Params · object | null

Model settings. For a managed voice model, exactly the control names its capabilities list; an unknown or out-of-range setting is a 422 unsupported_model_setting, never silently dropped.

idempotency_key
string | null

Keys the provider-job checkpoint: a retried call with the same key resumes the existing render instead of paying for a duplicate. On a managed voice model it keys the durable execution: the same key and input return the same task (replayed: true), a different input is 409 idempotency_conflict, and keys never expire.

Maximum string length: 200
wait
boolean
default:true

Managed voice models only. true (the default) runs the generation inside this call for up to about 210 seconds and returns the task envelope, terminal when it finished in time; false returns the queued task envelope at once — poll its poll_url no sooner than retry_after_ms.

quote
boolean
default:false

Managed voice models only. true validates the request and returns an audio_generation_quote (committed: false) with the usage estimate and a credit check. Nothing is started, held or charged, and the idempotency_key is ignored.

source_audio
string | null

speech_to_speech only: the recording to convert, as a file_ id from POST /v1/uploads (MP3, WAV or AAC).

source_audio_cleanup
enum<string> | null

speech_to_speech only. on_failure deletes a source recording this user uploaded for this generation if the generation fails or is canceled; never (the default) keeps it.

Available options:
on_failure,
never
context
GenerateAudioContext · object | null

Managed voice models with context stitching only: text or earlier tasks around this passage, so a regenerated passage matches its neighbours.

pronunciation_dictionaries
GenerateAudioPronunciationPin · object[] | null

Managed voice models whose pronunciation mode is dictionary: pinned dictionary revisions, in precedence order.

quote_source_seconds
number | null

speech_to_speech quotes without source_audio: the recording's length in seconds (above 0, at most 300).

max_estimated_credits
integer | null

Managed voice models only: the credit estimate the caller approved. A generation whose admission estimate is higher is refused (422 validation_failed, estimate_exceeds_approval) before anything is held or run.

Response

Successful Response

The response is of type Response Mediagenerateaudio · object.