ByteDance Launches Seed Audio 1.0 for Multi-Modal Sound Generation
2026-07-22 14:51

Woofun AI reports that ByteDance has released the Seed Audio 1.0 model, which generates dialogue, sound effects, and ambient audio simultaneously from a single prompt. The system supports text and reference audio inputs with 100ms timing precision, producing up to two minutes of consistent character voice per run.

Seed Audio 1.0 covers over 20 languages, enabling cross-language voice transfer and emotional tone adjustments. Official evaluations indicate an audio usability rate exceeding 90% in most scenarios, with naturalness MOS scores above 4 for the majority of supported languages. The model is currently accessible through the Volcano Ark Experience Center and offers API integration.

Disclaimer: Views are the author's own and do not represent the platform. Do not reproduce without permission. Content is for reference only, not investment advice. Trade at your own risk.
Tags:
Seed Audio 1.0
Volcano Ark Experience Center
ByteDance
Share:
back