
Seed Audio 1.0: An AI Generator for Complete Sound Scenes
Seed Audio 1.0 generates dialogue, ambience, music, and sound effects in one pass. Here is what it is for, how it differs from TTS, and how to try it.
Seed Audio 1.0 is an AI audio generator built for sound scenes, not isolated voice clips. You describe a place, the people in it, and what happens, then it produces a mixed clip that can include multi-speaker dialogue, emotion, ambience, background music, and foley-style effects.
That is a different job from ordinary text-to-speech. TTS reads a script. Seed Audio tries to stage the moment around the script.
What it generates
A typical scene prompt can ask for several layers at once:
- Two or more speakers, with tone, pacing, and accent
- Room tone, weather, traffic, crowds, or other ambience
- A music bed that matches the emotion
- Timed events such as footsteps, a door slam, or a whoosh
Each generation is a fully mixed, non-streaming clip, up to about two minutes. You can steer it with text, optional reference audio (up to three clips), or one reference image when the mood needs a visual cue.
The site’s own prompt formula is straightforward: scene + speaker + emotion + language + ambience + BGM + sound effects + timing.
