QwenAudio/qwen-audio-agent
A realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents
Explore 15 GitHub repositories focused on speech-recognition. Discover top-starred projects and those trending this week.
A realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
🧠 Leon is your open-source personal assistant.
Open-source, accurate and easy-to-use video speech recognition & clipping tool. LLM-based AI clipping integrated.
Production First and Production Ready End-to-End Speech Recognition Toolkit
Machine Learning and Agentic AI Resources, Practice and Research
这是一个全自动(音频)视频翻译项目。利用Whisper识别声音,AI大模型翻译字幕,最后合并字幕视频,生成翻译后的视频。
Transcribe any audio to text, translate and edit subtitles 100% locally with a web UI. Powered by whisper models!
🎙Speech recognition using the tensorflow deep learning framework, sequence-to-sequence neural networks
On-device, real-time multimodal AI. Have natural voice and vision conversations with an AI that runs entirely on your machine. Powered by Gemma 4 E2B and Kokoro.
SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
🧠 Leon is your open-source personal assistant.
Open-source, accurate and easy-to-use video speech recognition & clipping tool. LLM-based AI clipping integrated.
Production First and Production Ready End-to-End Speech Recognition Toolkit
Machine Learning and Agentic AI Resources, Practice and Research
这是一个全自动(音频)视频翻译项目。利用Whisper识别声音,AI大模型翻译字幕,最后合并字幕视频,生成翻译后的视频。
Transcribe any audio to text, translate and edit subtitles 100% locally with a web UI. Powered by whisper models!
A realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents
🎙Speech recognition using the tensorflow deep learning framework, sequence-to-sequence neural networks
On-device, real-time multimodal AI. Have natural voice and vision conversations with an AI that runs entirely on your machine. Powered by Gemma 4 E2B and Kokoro.
SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.
Free, open source voice dictation for macOS. On-device transcription with Apple's Speech framework. No cloud/no API keys/no account.
早耳 - Real-time multilingual speech-to-text on CPU only. Live subtitles, browser dashboard, speaker labels, translation. No GPU, no cloud.
An end-to-end framework for multi-speaker transcription that jointly models who spoke, when, and what.