Back to News
RSS feednews.ycombinator.com

Ask HN: Are There Models That Generate Sounds from Text and Reference Audio?

Summary

A Hacker News user is looking for a commercial AI model that can generate a new sound from both a reference recording and a text instruction. The use case is creating sound effects from a source with few clean examples, where text alone does not provide enough control over the result. The author notes that image systems commonly combine text prompts with reference images, while audio systems can already accept text and audio inputs for description or classification. They tried ElevenLabs SFX, but it supports text-to-audio without the desired reference input. They also tried Stable Audio 3, which is intended to support the workflow, but report that its results were poor. The post asks whether another solution exists and whether reference-guided audio generation is simply not yet a common commercial capability.