debpalash/OmniVoice-Studio

debpalash/OmniVoice-Studio — Local voice clone, video dubbing, dictation and audiobook maker. The open-source ElevenLabs alternative.

Updated 2026-07-28 · Source https://github.com/debpalash/OmniVoice-Studio · 中文版

What is the idea?

Local voice cloning, video dubbing, dictation and audiobook production in one open-source desktop app — the whole chain runs on your own machine.

Who would use it, and what do they get?

Strong, and it is the sanctioned exception — the output is a finished thing an ordinary person can hand to someone: an audiobook, a dubbed video.

Why would anyone bother?

No subscription, no per-character meter, audio never leaves your machine; 646 languages dwarf cloud tools; and it ships an MCP server, so Claude or Cursor can drive it directly — useful to humans and AI agents that need a local voice pipeline.

What is missing today?

Every commercial equivalent is a cloud subscription that requires uploading your voice, which is exactly the asset people are least willing to hand over — and the reason a huge fraction of would-be users never start.

What makes this one different?

Pick the one output that is unambiguously a finished artifact and unambiguously the user's own: my own book, read in my own voice, as a file I keep. Refuse third-party reference voices outright. Cold start CN = 豆瓣/小红书「有声书 自制」与独立作者群; EN = r/selfpublish audiobook-cost threads (ACX pricing complaints are a standing genre).

What is the sharpest thing about it?

Fully local, end-to-end dubbing on a hub of 14 TTS + 11 ASR engines — pick the best model per language/latency instead of being locked to one; auto-offloads on ≤8 GB VRAM and runs on pure CPU.

What could kill it?

Voice cloning is the single most abuse-prone capability in this catalogue; without hard consent gating on the reference voice, the product's most enthusiastic early users will be the wrong ones.

Who pays, and for what?

Buyer: dubbing shops and creators in long-tail languages ElevenLabs covers poorly, and studios that can't send client audio to a cloud. Money: a supported team edition + priced engine/model packs — the value is the workflow (batch, project files, invisible watermarking), not any single model.

What is the main risk?

Who kills it: ElevenLabs or Descript ship a solid offline tier, or one open model gets good enough that the 14-engine hub stops mattering. It's in active beta — reliability across 14 engines and three OSes is the hard part. Watch: does a real 30-min video dub cleanly in one click without hand-fixing timing.

What did real people actually say?

Verbatim from the public sources linked above, not paraphrased: “…d the repo terms both. This is the condition that most often ends a proposal, and the one I can't resolve for you. 2. **A `SubprocessBackend` sidecar.** Your instinct is right, and it's now a well-worn shape. `docs/engines/omni…” — github “Being straight with you: **nobody on the project has a Strix Halo / gfx1151 machine**, so I can't verify a fix for this — and shipping an unverified ROCm change risks breaking the AMD users who currently work. Here's what the c…” — github

How much work is it?

about two months of work

Where does the evidence come from?

https://github.com/debpalash/OmniVoice-Studio

GitHubWorth buildingTwo months of workHas verbatim voices

SoundGate Guitar  ·  Estera