Open-source repositories gaining traction right now.
6 repositories
#speech-to-text
Build local voice agents with open-source models

Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.

Open Source Voice Agent Platform

Local AI anywhere, for everyone — LLM inference, chat UI, voice, agents, workflows, RAG, and image generation. No cloud, no subscriptions.

Real-time speech-to-text caption appliance for a deaf user. Raspberry Pi + 10" touchscreen that transcribes phone calls and room conversation in near real-time.

Push-to-talk voice typing for your terminal. Local Whisper, cross-platform.
Paste a github.com URL. Submissions are reviewed before they are tracked.