πŸŽ™οΈ SenseVoice

Speech Recognition + Emotion Detection + Audio Events β€” All in One Model

5 languages (zh/en/yue/ja/ko) Β· 7x faster than Whisper-small Β· 17x faster than Whisper-large

⭐ GitHub Β· πŸ› οΈ FunASR Toolkit Β· πŸ“„ Paper Β· πŸš€ Fun-ASR (31 Languages)

How it works: Upload audio or record via microphone β†’ SenseVoice transcribes speech and detects emotions (πŸ˜ŠπŸ˜‘πŸ˜”) and sound events (πŸŽΌπŸ‘πŸ˜€πŸ˜­πŸ€§). Event labels appear at the front of text, emotions at the end.
Language
Examples
Upload audio or use microphone Language