Projects with this topic
-
Transcribes audio files to text locally with WhisperX.
Local audio-to-text web app. Upload an audio file (MP3/WAV/M4A), a background job transcribes it with WhisperX, then view and download the transcript as JSON, SRT, TXT, or Markdown.
Spring Boot + Thymeleaf frontend, Python worker under the hood.
Runs locally, no cloud.
Updated -
A robust CLI that turns recordings and PDFs into clean, structured Markdown notes end-to-end.
Audio — chunks with configurable length and overlap, transcribes via a dedicated Whisper STT endpoint or cascades to OpenAI-compatible multimodal and Native Ollama backends.
PDF — text PDFs pass through pdftotext straight to refinement; scanned PDFs are auto-detected by character count, rendered to page images, and read by a vision model (or tesseract via an explicit --ocr escape, never silently).
LLM refinement — dedupes overlapping boundaries and formats the raw transcript into clean notes.
Resumable content-addressed cache — keyed by input hash with per-chunk/page signatures; a changed parameter invalidates only what depends on it.
Config resolves CLI > env var > settings file > defaults, including persistent chunk sizes. Python 3.12, built with uv, 147 tests at 100% coverage. Architecture is protocol-based and documented via TIPs (docs/).
Updated -
Audio capture service for the Oremi smart home ecosystem with wake word detection, speech transcription, and Home Assistant integration via MQTT.
Updated -
Transcripción local de alta fidelidad en Apple Silicon (Metal GPU con MLX). Soporte multipista para grabaciones de OBS (.mka), mezcla de audio y diálogo entrelazado en Markdown.
Updated -
Self-hosted speech server in one Docker image. OpenAI-compatible /v1/audio/transcriptions and /v1/audio/speech across 12 ASR models (Whisper, Parakeet, Canary, Sherpa-ONNX, Vosk) and 3 TTS engines (Kokoro, Qwen3-TTS voice cloning, Chatterbox Turbo). Live WebSocket ASR, file staging, MCP built in. CPU + CUDA images.
Updated -
A second brain for conversations you're already having: listens to your mic and system audio as two separate streams, transcribes both live, works out when the other side shut up, and hands you a reply draft before you panic. It drafts, never sends or speaks for you. Alpha — live capture is Linux/PipeWire only.
Updated -
Apply LLM prompts to a transcript in consistent ways.
Updated -
Fonction JavaScript pour créer des rébus avec des emojis à partir de phrase en français. (peut aussi faire de la transcription phonétique)
Mirror of https://github.com/ptlc8/rebus
Updated -
Drop a file, get a transcript, nobody else ever hears it. Local Whisper + pyannote transcription with speaker labels, on your own machine, Windows and Linux.
Updated -
System-wide voice-to-text for macOS. Hold a hotkey, speak, text appears at your cursor. Supports OpenAI, Gemini, Apple Speech & local models. No subscription — bring your own API keys.
Updated -
A Node.js application for transcribing audio files into word-level and phrase-level subtitles in SRT format.
Updated -
Transcripty is a privacy-first, locally run transcription software using locally running AI-models for speech to text analysis. The software has light editing capability for splitting and merging text blocks and assigning speakers. The transcription result can be exported to a text doc for use in other applications.
Updated -
MacWhisperer is a macOS desktop GUI for OpenAI Whisper that makes local speech transcription and English translation easy, with auto model loading, live recording, input device controls, and file-based workflows.
Updated -
(Design WIP) Ext. tool adding a transcription (OCR) workflow to the EmuHawk (BizHawk) emulator, allowing retro games to be translated partially- or fully-automatically
Updated -
A simple phonemic transcriber for the american english language.
Updated -
Un simple transcriptor fonológico para la lengua española.
Updated -
A simple universal transcriber for languages with unicode characters.
Updated -
Groq AI Transcribe is a simple Python-based project using the Groq API to transcribe audio files. The project provides a simple and efficient way to transcribe audio files using the Groq API via the command line.
Updated -
web components that record audio and transcribe it to text using openai api's https://attention1.gitlab.io/ai-interface
Updated