Stable Audio 3 Small Music
Connect Claude, Cursor, VS Code, or Codex to Layer’s MCP server and Stable Audio 3 Small Music becomes a tool your agent can reach for — it picks the parameters, runs the generation, and shows you the result in the conversation.
https://mcp.app.layer.ai/mcpAsk your agent
“Use Stable Audio 3 Small Music for a heavy stone door grinding open in a dungeon.”
execute_forge({
"base_model_id": "stable-audio-3-small-music",
"prompt": "A heavy stone door grinding open in a dungeon",
"batch_size": 1
})Authenticate with a personal access token, post the model’s own form, and poll the inference until it completes. Every field below is accepted by this endpoint.
# Start the generation — returns an inference id straight away.
curl -X POST https://api.app.layer.ai/api/v2/workspaces/$WORKSPACE_ID/base-models/stable-audio-3-small-music/inferences \
-H "Authorization: Bearer $LAYER_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"prompt": "A heavy stone door grinding open in a dungeon"
}'
# Poll until status is "complete" — the response then carries the asset URLs.
curl https://api.app.layer.ai/api/v2/workspaces/$WORKSPACE_ID/inferences/$INFERENCE_ID \
-H "Authorization: Bearer $LAYER_TOKEN"- prompt string
- Describe what to generate, or the change to make.
- negative_prompt string
- What to keep out of the result.
- prompt_enhancement enumdefault AUTOAUTO, OFF, STANDARD
- Let the model rewrite the prompt before generating.
- seed integer
- The starting point for the model's randomness. The same seed and settings give the same result.
- guidance_scale numberdefault 10–25
- How closely the result follows the prompt. Lower gives the model more freedom.
- inference_steps integerdefault 81–100
- How many refinement passes the model makes. More add detail and take longer.
- duration_seconds numberdefault 301–120
- Length of the generated clip, in seconds.
- guidance_file_init_audio file[]
- The audio to modify
- extend_seconds number1–120
- How much new audio continues the clip past its end.
- audio_sections array
- The stretches of the clip the prompt rewrites; the rest is kept.
- prompt_language string
- The language the prompt is written in, so it reaches the model as you meant it.
Why use Stable Audio 3 Small Music for audio?
Natural voice generation
Stable Audio 3 Small Music produces expressive, high-quality audio suitable for character dialogue, narration, and voiceover work.
Flexible content types
Supports a range of use cases from in-game dialogue and cinematics to marketing narration and social media content.
Fast iteration
Generate and refine audio content quickly, enabling rapid prototyping of character voices and sound design.
Character voiceover and dialogue production
Generate expressive character voices for in-game dialogue, cutscenes, and interactive narratives. Iterate on tone and delivery rapidly.
Marketing narration and promotional audio
Create professional voiceovers for trailers, app store videos, and social media content without booking voice talent.
Sound design exploration and prototyping
Quickly prototype sound effects, ambient audio, and musical elements to test creative directions early in production.
FAQ
How much does Stable Audio 3 Small Music cost? Is it free?+
How does Layer's pricing work?+
There are no seat fees, feature gates, or plans on Layer. Instead, our platform uses a consumption based system with Creative Units (CUs) with a flexible monthly subscription. Every generation on Layer (image, video, 3D, or audio) consumes a Creative Unit, and you only pay for what you create.
Other audio models on Layer
Stable Audio 2.5
Stability AI
Enterprise-grade audio generation with multi-part compositions and audio inpainting.
Stable Audio 3 Small SFX
Stability AI
Fast sound effects with restyle, extend and inpaint.
Stable Audio 3 Medium
Stability AI
Full songs and long scores up to six minutes, with restyle, extend and inpaint.
VEED Clean Audio 1.0
VEED
Remove background noise from speech while keeping quiet and distant words intact.
Mureka 9.5 Song
Mureka
Complete songs with vocals from a description or supplied lyrics.
ElevenLabs TTS V4 Turbo
ElevenLabs
Lower-latency Eleven v4 Turbo speech with audio tags, stability, and similarity control.