Login
Sign Up
Woofun AI reports that ByteDance has released the Seed Audio 1.0 model, which generates dialogue, sound effects, and ambient audio simultaneously from a single prompt. The system supports text and reference audio inputs with 100ms timing precision, producing up to two minutes of consistent character voice per run.
Seed Audio 1.0 covers over 20 languages, enabling cross-language voice transfer and emotional tone adjustments. Official evaluations indicate an audio usability rate exceeding 90% in most scenarios, with naturalness MOS scores above 4 for the majority of supported languages. The model is currently accessible through the Volcano Ark Experience Center and offers API integration.