stt
Tracked open-source repos tagged stt, sorted by stars.
Related topics
Topics that frequently appear alongside stt on the same repo.
Recent risers
Repos created in the last 90 days, tagged stt.
- #1
The spend meter and budget gate for AI voice agents. Meters STT + TTS + LLM + telephony per call, out of the box (Pipecat, LiveKit — Python & TypeScript). Hard-stops the next turn before it crosses your ceiling. Local, no account, no telemetry. Built by Floe — cost controls for Voice AI.
★ 434
- #1
Your AI second brain. Self-hostable. Get answers from the web or your docs. Build custom agents, schedule automations, do deep research. Turn any online or local LLM into your personal, autonomous AI (gpt, claude, gemini, llama, qwen, mistral). Get started - free.
★ 37,010+223Star change over the last 7 days - #2
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
★ 15,101+17Star change over the last 7 days - #3
Very low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces
★ 10,997+33Star change over the last 7 days - #4
Open Vision Agents by Stream. Build voice and vision agents quickly with any model or video provider. Uses Stream's edge network for ultra-low latency.
★ 8,113+9Star change over the last 7 days - #5★ 4,780+9Star change over the last 7 days
- #6
Transcribe any audio to text, translate and edit subtitles 100% locally with a web UI. Powered by whisper models!
★ 3,068+2Star change over the last 7 days - #7
🐸STT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy.
★ 2,606+1Star change over the last 7 days - #8
🎙Speech recognition using the tensorflow deep learning framework, sequence-to-sequence neural networks
★ 2,173+0Star change over the last 7 days - #9
Realtime Voice AI with 100+ Models on Arduino ESP32 with Secure Websockets and Edge Functions for AI Companions, and Devices
★ 1,938+5Star change over the last 7 days - #10
Meet Ava, the WhatsApp Agent
★ 1,675+2Star change over the last 7 days - #11
Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation.
★ 1,622+8Star change over the last 7 days - #12
Dicio assistant app for Android
★ 1,463+3Star change over the last 7 days - #13
🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.
★ 1,418+0Star change over the last 7 days - #14
Synchronized Translation for Videos. Video dubbing
★ 1,413+2Star change over the last 7 days - #15
💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
★ 1,399+0Star change over the last 7 days - #16
小智ESP32的Java企业级管理平台,提供设备监控、音色定制、角色切换和对话记录管理的前后端及服务端一体化解决方案
★ 1,337+1Star change over the last 7 days - #17
Gp.nvim (GPT prompt) Neovim AI plugin: ChatGPT sessions & Instructable text/code operations & Speech to text [OpenAI, Ollama, Anthropic, ..]
★ 1,320+0Star change over the last 7 days - #18
Real-time speech translation — macOS & Windows, free TTS, no server, your API keys only
★ 1,278+2Star change over the last 7 days - #19
Real-time two-way speech translation for bilingual meetings — auto-detects the spoken language and translates both directions, cloud or fully offline on-device. Desktop (Windows · macOS · Linux) + browser extension (Chrome · Edge) for Zoom, Meet, Teams & any app.
★ 1,186+40Star change over the last 7 days - #20
NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.
★ 1,094+13Star change over the last 7 days - #21★ 805+1Star change over the last 7 days
- #22
🎙️ An intelligent voice productivity assistant that turns speech into clean text, useful actions, and structured knowledge. It helps users capture ideas, communicate naturally, automate repetitive tasks, and stay productive across different apps and workflows.
★ 804+3Star change over the last 7 days - #23
Speech to Text to Speech. Song now playing. Sends text as OSC messages to VRChat to display on avatar. (STTTS) (Speech to TTS) (VRC STT System) (VTuber TTS)
★ 802+3Star change over the last 7 days - #24
A modular Swift SDK for audio processing with MLX on Apple Silicon
★ 772+6Star change over the last 7 days - #25
Running speech to text model (whisper.cpp) in Unity3d on your local machine.
★ 751+1Star change over the last 7 days - #26
MLX Omni Server is a local inference server powered by Apple's MLX framework, specifically designed for Apple Silicon (M-series) chips. It implements OpenAI-compatible API endpoints, enabling seamless integration with existing OpenAI SDK clients while leveraging the power of local ML inference.
★ 741+1Star change over the last 7 days - #27★ 671+0Star change over the last 7 days
- #28
:speech_balloon: /so.nus/ STT (speech to text) for Node with offline hotword detection
★ 638+0Star change over the last 7 days - #29
A React component to make correcting automated transcriptions of audio and video easier and faster. By BBC News Labs. - Work in progress
★ 620-1Star change over the last 7 days - #30
C++ ggml runtime hub for multilingual ASR and TTS models: Cohere Transcribe, Parakeet TDT, Voxtral, Canary 1B v2, etc, plus universal forced alignment, and more
★ 610+21Star change over the last 7 days