Skip to main content
Quail Voice Focus isolates the primary speaker. It is designed for settings where one or more speakers may be present but the application targets a single speaker, suppressing interfering speech, residual echo, and background noise. Voice Focus listens for a moment at the start of each session before applying suppression. During this warm-up, audio may sound closer to the original. Once a clear primary speaker is detected (typically within a few seconds) full suppression kicks in. The sooner the primary speaker talks, the shorter the warm-up period. On very short clips where the primary speaker doesn’t get a chance to speak, suppression may not fully activate. This is by design, as Voice Focus prioritizes accuracy over speed and will not suppress a speaker it hasn’t confidently identified yet. Quail Voice Focus may also be used as a pre-processing step for third-party VADs that do not perform well in noisy and multi-speaker conditions, as it will suppress background noise and interfering speech.
See Improving ASR with Voice Focus and the model reference for available sizes and specs.