modelscope/FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Explore 7 GitHub repositories focused on asr. Discover top-starred projects and those trending this week.
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Production First and Production Ready End-to-End Speech Recognition Toolkit
SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.
First foundation ASR built for the real world - 7 atomic acoustic conditions, 54 compound scenarios, 2.6M samples, and up to ~30% gains over SOTA where every other model falls apart. **You'll come back to MEGA-ASR, after the rest fail in the wild. ⭐**
🎙️Voice input and translation app for macOS. Press to talk, release to paste.
早耳 - Real-time multilingual speech-to-text on CPU only. Live subtitles, browser dashboard, speaker labels, translation. No GPU, no cloud.
An end-to-end framework for multi-speaker transcription that jointly models who spoke, when, and what.
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Production First and Production Ready End-to-End Speech Recognition Toolkit
SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.
First foundation ASR built for the real world - 7 atomic acoustic conditions, 54 compound scenarios, 2.6M samples, and up to ~30% gains over SOTA where every other model falls apart. **You'll come back to MEGA-ASR, after the rest fail in the wild. ⭐**
🎙️Voice input and translation app for macOS. Press to talk, release to paste.
早耳 - Real-time multilingual speech-to-text on CPU only. Live subtitles, browser dashboard, speaker labels, translation. No GPU, no cloud.
An end-to-end framework for multi-speaker transcription that jointly models who spoke, when, and what.