debpalash/OmniVoice-Studio
debpalash/OmniVoice-Studio — Local voice clone, video dubbing, dictation and audiobook maker. The open-source ElevenLabs alternative.
What is the idea?
Local voice cloning, video dubbing, dictation and audiobook production in one open-source desktop app — the whole chain runs on your own machine.
Who would use it, and what do they get?
Strong, and it is the sanctioned exception — the output is a finished thing an ordinary person can hand to someone: an audiobook, a dubbed video.
Why would anyone bother?
No subscription, no per-character meter, audio never leaves your machine; 646 languages dwarf cloud tools; and it ships an MCP server, so Claude or Cursor can drive it directly — useful to humans and AI agents that need a local voice pipeline.
What is missing today?
Every commercial equivalent is a cloud subscription that requires uploading your voice, which is exactly the asset people are least willing to hand over — and the reason a huge fraction of would-be users never start.
What makes this one different?
Pick the one output that is unambiguously a finished artifact and unambiguously the user's own: my own book, read in my own voice, as a file I keep. Refuse third-party reference voices outright. Cold start CN = 豆瓣/小红书「有声书 自制」与独立作者群; EN = r/selfpublish audiobook-cost threads (ACX pricing complaints are a standing genre).
What is the sharpest thing about it?
Fully local, end-to-end dubbing on a hub of 14 TTS + 11 ASR engines — pick the best model per language/latency instead of being locked to one; auto-offloads on ≤8 GB VRAM and runs on pure CPU.
What could kill it?
Voice cloning is the single most abuse-prone capability in this catalogue; without hard consent gating on the reference voice, the product's most enthusiastic early users will be the wrong ones.
Who pays, and for what?
Buyer: dubbing shops and creators in long-tail languages ElevenLabs covers poorly, and studios that can't send client audio to a cloud. Money: a supported team edition + priced engine/model packs — the value is the workflow (batch, project files, invisible watermarking), not any single model.
What is the main risk?
Who kills it: ElevenLabs or Descript ship a solid offline tier, or one open model gets good enough that the 14-engine hub stops mattering. It's in active beta — reliability across 14 engines and three OSes is the hard part. Watch: does a real 30-min video dub cleanly in one click without hand-fixing timing.
What did real people actually say?
Verbatim from the public sources linked above, not paraphrased: “…d the repo terms both. This is the condition that most often ends a proposal, and the one I can't resolve for you. 2. **A `SubprocessBackend` sidecar.** Your instinct is right, and it's now a well-worn shape. `docs/engines/omni…” — github “Being straight with you: **nobody on the project has a Strix Halo / gfx1151 machine**, so I can't verify a fix for this — and shipping an unverified ROCm change risks breaking the AMD users who currently work. Here's what the c…” — github
How much work is it?
about two months of work
Where does the evidence come from?
https://github.com/debpalash/OmniVoice-Studio
← SoundGate Guitar · Estera →