Projects with this topic
-
Transcribes audio files to text locally with WhisperX.
Local audio-to-text web app. Upload an audio file (MP3/WAV/M4A), a background job transcribes it with WhisperX, then view and download the transcript as JSON, SRT, TXT, or Markdown.
Spring Boot + Thymeleaf frontend, Python worker under the hood.
Runs locally, no cloud.
Updated -
Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation
Updated -
-
Self-hosted speech server in one Docker image. OpenAI-compatible /v1/audio/transcriptions and /v1/audio/speech across 12 ASR models (Whisper, Parakeet, Canary, Sherpa-ONNX, Vosk) and 3 TTS engines (Kokoro, Qwen3-TTS voice cloning, Chatterbox Turbo). Live WebSocket ASR, file staging, MCP built in. CPU + CUDA images.
Updated -
A second brain for conversations you're already having: listens to your mic and system audio as two separate streams, transcribes both live, works out when the other side shut up, and hands you a reply draft before you panic. It drafts, never sends or speaks for you. Alpha — live capture is Linux/PipeWire only.
Updated -
Give your AI agent a voice. Local speech stack for Apple Silicon: TTS, ASR, forced alignment, voices, daemon, and MCP bridge.
Updated