GitStar
GitHub TrendingTopicsLanguages
/

Speech To Text GitHub Repositories

Explore popular GitHub repositories tagged “speech-to-text”.

Compare stars, forks, and programming language using the same GitStar view as GitHub Trending.

RepositoriesGitHub topic
Ranked by:Stars

Trending Repositories

ggml-org/whisper.cpp

Port of OpenAI's Whisper model in C/C++

C++52,9846,076
cjpais/Handy

A free, open source, and extensible speech-to-text application that works completely offline.

Rust29,8502,648
Zackriya-Solutions/meetily

Privacy first, AI meeting assistant with 4x faster Parakeet/Whisper live transcription, speaker diarization, and Ollama summarization built on Rust. 100% local processing. no cloud required. Meetily (Meetly Ai - https://meetily.ai) is the #1 Self-hosted, Open-source Ai meeting note taker for macOS & Windows. Understand How to write meeting minutes

Rust29,3523,126
mozilla-ai/llamafile

Distribute and run LLMs with a single file.

C++25,6251,558
SYSTRAN/faster-whisper

Faster Whisper transcription with CTranslate2

Python24,9672,028
m-bain/whisperX

WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)

Python23,6212,387
screenpipe/screenpipe

YC (S26) | Record your screen 24/7 and plug into your agents. Local, private, secure. Connect to OpenClaw, Hermes agent and 100+ apps

Rust21,0472,107
modelscope/FunASR

Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.

Python19,9051,990
jianchang512/pyvideotrans

Translate the video from one language to another and embed dubbing & subtitles.

Python18,7222,316
leon-ai/leon

🧠 Leon is your open-source personal assistant.

TypeScript17,4421,463
kaldi-asr/kaldi

kaldi-asr/kaldi is the official location of the Kaldi project.

Shell15,4595,355
alphacep/vosk-api

Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node

Jupyter Notebook15,0641,744
k2-fsa/sherpa-onnx

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages

C++14,2371,632
huggingface/speech-to-speech

Build local voice agents with open-source models

Python12,5981,544
abus-aikorea/voice-pro

Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.

Python12,4521,802
speechbrain/speechbrain

A PyTorch-based Speech Toolkit

Python11,7601,717
QuentinFuxa/WhisperLiveKit

Real-time, local speech-to-text with streaming ASR, speaker diarization, translation, and OpenAI/Deepgram-compatible APIs.

Python10,6151,103
KoljaB/RealtimeSTT

A robust, efficient, low-latency speech-to-text library with advanced voice activity detection, wake word activation and instant transcription.

Python10,058848
debpalash/VoiceStudio

VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages. No accounts, no API keys, no cloud.

Python10,0551,670
QwenAudio/SenseVoice

Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.

C9,101808
Uberi/speech_recognition

Speech recognition module for Python, supporting several engines and APIs, online and offline.

Python8,9832,417
nl8590687/ASRT_SpeechRecognition

A Deep-Learning-Based Chinese Speech Recognition System 基于深度学习的中文语音识别系统

Python8,3811,894
Blaizzy/mlx-audio

A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.

Python7,754690
TalAter/annyang

💬 Speech recognition for your site

TypeScript6,8161,048
argmaxinc/argmax-oss-swift

On-device Speech AI for Apple Silicon

Swift6,328598
modelscope/FunClip

FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.

Python6,160737
snakers4/silero-models

Silero Models: pre-trained text-to-speech models made embarrassingly simple

Jupyter Notebook6,067372
MahmoudAshraf97/whisper-diarization

Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper

Jupyter Notebook5,626502
OpenWhispr/openwhispr

Voice-to-text dictation app with local (Nvidia Parakeet/Whisper) and cloud models (BYOK). Privacy-first and available cross-platform.

JavaScript5,532780
dograh-hq/dograh

Open source voice AI platform. Self-hosted alternative to Vapi and Retell. On Prem, BYOK across Speech to Speech or LLM/STT/TTS, with a visual workflow builder, MCP native and telephony support.

Python5,4001,310
GitStar

See what the GitStar community is most excited about today.

Trending

GitHub Trending TodayGitHub Trending WeeklyGitHub Trending Monthly

Languages

Browse all languagesTrending PythonTrending JavaScript

Explore

Browse GitHub topicsAI repositoriesDeveloper tools
© 2026 GitStarGitHub Trending source