ElevenLabs Music v2.
Generates a song from a text prompt or a chunk-based composition plan
with the music_v2 model: better prompt adherence, composition and
vocal delivery than v1, plus multilingual output.
The client provides a convenient way to interact with the model API.
npm install --save @fal-ai/clientThe @fal-ai/serverless-client package has been deprecated in favor of @fal-ai/client. Install the new package and update your imports — see client setup.
Set FAL_KEY as an environment variable in your runtime.
export FAL_KEY="YOUR_API_KEY"The client API handles the API submit protocol. It will handle the request status updates and return the result when the request is completed.
import { fal } from "@fal-ai/client";
const result = await fal.subscribe("elevenlabs/music/v2", {
input: {},
logs: true,
onQueueUpdate: (update) => {
if (update.status === "IN_PROGRESS") {
update.logs.map((log) => log.message).forEach(console.log);
}
},
});
console.log(result.data);
console.log(result.requestId);The API uses an API Key for authentication. It is recommended you set the FAL_KEY environment variable in your runtime when possible.
import { fal } from "@fal-ai/client";
fal.config({
credentials: "YOUR_FAL_KEY"
});When running code on the client-side (e.g. in a browser, mobile app or GUI applications), make sure to not expose your FAL_KEY. Instead, use a server-side proxy to make requests to the API. For more information, check out our server-side integration guide.
The client API provides a convenient way to submit requests to the model.
import { fal } from "@fal-ai/client";
const { request_id } = await fal.queue.submit("elevenlabs/music/v2", {
input: {},
webhookUrl: "https://optional.webhook.url/for/results",
});You can fetch the status of a request to check if it is completed or still in progress.
import { fal } from "@fal-ai/client";
const status = await fal.queue.status("elevenlabs/music/v2", {
requestId: "764cabcf-b745-4b3e-ae38-1200304cf45b",
logs: true,
});Once the request is completed, you can fetch the result. See the Output Schema for the expected result format.
import { fal } from "@fal-ai/client";
const result = await fal.queue.result("elevenlabs/music/v2", {
requestId: "764cabcf-b745-4b3e-ae38-1200304cf45b"
});
console.log(result.data);
console.log(result.requestId);Some attributes in the API accept file URLs as input. Whenever that's the case you can pass your own URL or a Base64 data URI.
You can pass a Base64 data URI as a file input. The API will handle the file decoding for you. Keep in mind that for large files, this alternative although convenient can impact the request performance.
You can also pass your own URLs as long as they are publicly accessible. Be aware that some hosts might block cross-site requests, rate-limit, or consider the request as a bot.
We provide a convenient file storage that allows you to upload files and use them in your requests. You can upload files using the client API and use the returned URL in your requests.
import { fal } from "@fal-ai/client";
const file = new File(["Hello, World!"], "hello.txt", { type: "text/plain" });
const url = await fal.storage.upload(file);The client will auto-upload the file for you if you pass a binary object (e.g. File, Data).
Read more about file handling in our file upload guide.
prompt stringThe text prompt describing the music to generate
The chunk-based composition plan for the music
music_length_ms integerThe length of the song to generate in milliseconds. Used only in conjunction with prompt. Must be between 3000ms and 600000ms. Optional - if not provided, the model will choose a length based on the prompt.
force_instrumental booleanIf true, guarantees that the generated song will be instrumental. If false, the song may or may not be instrumental depending on the prompt. Can only be used with prompt.
seed integerRandom seed to initialize the music generation process. Can only be used with composition_plan. The same seed with the same parameters gives more consistent results, but exact reproducibility is not guaranteed.
output_format OutputFormatEnumOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. Note that the μ-law format (sometimes written mu-law, often approximated as u-law) is commonly used for Twilio audio inputs. Default value: "mp3_48000_192"
Possible enum values: mp3_22050_32, mp3_24000_48, mp3_44100_32, mp3_44100_64, mp3_44100_96, mp3_44100_128, mp3_44100_192, mp3_48000_128, mp3_48000_192, mp3_48000_240, mp3_48000_320, pcm_8000, pcm_16000, pcm_22050, pcm_24000, pcm_32000, pcm_44100, pcm_48000, ulaw_8000, alaw_8000, opus_48000_32, opus_48000_64, opus_48000_96, opus_48000_128, opus_48000_192
{
"prompt": "Mysterious original soundtrack, themes of jungle, rainforest, nature, woodwinds, busy rhythmic tribal percussion.",
"output_format": "mp3_48000_192"
}The generated music audio file in MP3 format
{
"audio": {
"content_type": "audio/mpeg",
"url": "https://storage.googleapis.com/falserverless/example_outputs/elevenlabs/music_generated.mp3",
"file_name": "music_generated.mp3"
}
}The chunks that make up the generation, in order. At most 30 chunks, totalling between 3000ms and 600000ms.
url string* requiredThe URL where the file can be downloaded from.
content_type stringThe mime type of the file.
file_name stringThe name of the file. It will be auto-generated if not provided.
file_size integerThe size of the file in bytes.
text string* requiredThe text to generate for this chunk. Can start with an optional section name in square brackets, e.g. [Verse 1], followed by lyric lines, plus inline directions in curly braces, e.g. {scratching}. Section names must be between 1 and 100 characters. At most 30 lines are allowed, each at most 200 characters.
duration_ms integer* requiredThe duration of the chunk in milliseconds. Must be between 3000ms and 120000ms.
The styles and musical directions that should be present in this chunk. Use English for best results. The styles of the first chunk matter most as they set the overall tone and genre; later chunks can add nuance, progression or change direction.
The styles and musical directions that should not be present in this chunk. Leaving this empty is a good default; only set it to explicitly avoid a particular style or direction.
context_adherence EnumHow closely this chunk follows the context of its surrounding chunks. Low adherence lets the model deviate and be more creative, high adherence keeps it consistent with the surrounding context.
Possible enum values: low, medium, high
An audio clip whose style conditions this chunk. The reference on the first chunk matters most: it influences every later chunk, so condition from the first chunk to style the whole song.
audio_url string* requiredURL of the audio clip to use as the style reference.
start_ms integerOffset into the clip where the referenced window starts, in milliseconds.
end_ms integerOffset into the clip where the referenced window ends, in milliseconds. Defaults to 30000ms after start_ms, or the end of the clip if it is shorter. The window must be at most 30000ms long.
strength EnumHow strongly the model follows the reference. Low lets it deviate and be more creative, high keeps it close to the reference.
Possible enum values: low, medium, high, xhigh