# SAM Audio Separate

> Audio AI Model by Meta — available on Layer

Foundation model for audio source separation using text, visual, or temporal prompts.

SAM Audio (Segment Anything in Audio) from Meta AI is a foundation model for general audio separation that unifies text, visual, and temporal span prompting in a single framework. Built on a diffusion transformer architecture trained with flow matching on large-scale audio data, it isolates specific sounds from a mixture based on a natural language description — returning both the isolated target and the residual. Supports speech, music, and general sound separation.

## Specifications

| Property | Value |
|----------|-------|
| Provider | Meta |
| Category | Audio |
| Price | 1.4 / generation per generation |

## Capabilities

- Auto Duration

---

[Try SAM Audio Separate](https://next.app.layer.ai/me/session/new?baseModelId=SAM_AUDIO_SEPARATE) | [All Audio Models](https://layer.ai/models/audio) | [All Models](https://layer.ai/models)
