Projects with this topic
-
Self-hosted speech server in one Docker image. OpenAI-compatible /v1/audio/transcriptions and /v1/audio/speech across 12 ASR models (Whisper, Parakeet, Canary, Sherpa-ONNX, Vosk) and 3 TTS engines (Kokoro, Qwen3-TTS voice cloning, Chatterbox Turbo). Live WebSocket ASR, file staging, MCP built in. CPU + CUDA images.
Updated -
Containerized media processing tools over SSH. Drop files in, run ffmpeg/sox/imagemagick over SSH, get your shit out. No shell access, no bullshit - just a locked-down Python wrapper that only lets you run what you're supposed to run.
Updated -
Audiovista is an open-source cinematic AI engine that synchronizes audio and video with professional precision, including dialogue, sound effects, ambient audio, and adaptive music. Physics-aware, predictive, and interactive, it supports real-time previews, multi-track editing, VR/AR output, and cross-media export, empowering creators to produce immersive, AI-assisted cinematic experiences. https://roxanneardary.com/audiovista/
Updated -
-
Convert WAV audio files to a compressed audio format using the Web Audio API and MediaRecorder, enti
Updated -
Convert MP3 audio files to uncompressed WAV format using the Web Audio API, entirely in the browser.
Updated -
Trim audio files to a precise time range with waveform visualization, entirely in the browser.
Updated -
[2025] Raspberry Pi control system for interactive art installations. Orchestrates LED lighting, FM radio, live audio mixing, and environmental monitoring through unified Flask interface with 24/7 exhibition reliability and automatic service management.
Updated