Tools/debpalash/VoiceStudio

VoiceStudio

The open-source ElevenLabs alternative for local voice cloning, design, create, dubbing and dictation Desktop App

50.8k ★+15.5k/wkestablishedPythonAGPL-3.0trending

The Lens

By Erik Loyd, SaaS CEO and former COO/CFO of an AWS Premier Partner.

Updated Sep 2026

VoiceStudio (formerly OmniVoice Studio) is a local, open source answer to ElevenLabs. It clones and designs voices, does text-to-speech across 16 engines and transcription across 11, dubs video by transcribing, translating and re-voicing it, builds audiobooks, and runs a system-wide dictation widget. The app is AGPL-3.0, free, and runs on your own machine.

A GPU is optional but makes the difference. Minimums are 8 GB of RAM and 10 GB of disk, and the default workflow wants 8 GB or more of VRAM on NVIDIA or Apple Silicon. ROCm is Linux-only and opt-in. Installers cover Apple Silicon Macs, Windows 10/11 and recent Linux; Docker images are amd64 only, and Intel Macs need a remote backend.

Solo creators and small teams doing voiceover, dubbing or audiobooks get no per-character billing and no audio leaving the machine. It also exposes a local REST API, an OpenAI-compatible audio API and an MCP server. Companies embedding it in a closed product can get a commercial license by enquiry, with no published price, and a hosted Cloud API is in early access.

The catch is the model licenses, not the app's. The default voice engine's pretrained weights are CC-BY-NC, noncommercial, so paid voiceover work means switching to an Apache-2.0 engine like CosyVoice 3 or VoxCPM2. It is still labeled active beta. Watermarking is on by default, but responsible voice cloning is on you.

Free vs Self-Hosted vs Paid

fully free

Free tier: The app is AGPL-3.0-only with no paid tier: voice cloning and design, 16 TTS and 11 ASR engines, video dubbing, audiobooks with EPUB and PDF import, dictation, AudioSeal watermarking, a local REST/SSE/WebSocket API, an OpenAI-compatible audio API and an MCP server. Model weights carry their own licenses: the default OmniVoice weights are CC-BY-NC (noncommercial), while CosyVoice 3, VoxCPM2 and MOSS-TTS-Nano are Apache-2.0.

Self-hosted: Runs locally. Minimum 8 GB RAM, 10 GB free disk, and 4 GB VRAM for GPU acceleration; recommended 16 GB+ RAM, a 20 GB+ SSD and 8 GB+ VRAM, with large optional engines needing more. NVIDIA CUDA or Apple Silicon recommended; ROCm is Linux-only and opt-in, and Windows AMD machines fall back to CPU. Desktop builds for macOS 13.3+ on Apple Silicon, Windows 10/11 x64 and Linux x86_64 with glibc 2.39+. Docker images are linux/amd64 only. Intel Macs cannot run the local backend.

Paid: Nothing to buy today. A commercial license for embedding VoiceStudio in closed-source products is available by enquiry, with pricing tiers listed as coming soon. A hosted Cloud API is in early access with free usage credits.

Free under AGPL-3.0, but the default voice model's weights are noncommercial. Commercial embedding licenses are enquiry-only.

What to do by team size

Solo
free; pick an Apache-2.0 engine for paid work
Small team
free, if you check each engine's model license
Medium team
free for internal use; AGPL applies if you host it for others
Large team
Commercial license by enquiry if you embed it in a closed product
Self-hosting ops:moderate

Get tools like this every Wednesday

One featured tool, three on the radar. No fluff.

Similar Tools

Score
81/100 · A
Adoption30/30
Maintenance25/25
Community9/20
License7/15
Analysis10/10

A low score is not a verdict on quality. Young and niche tools start low by design. How we calculate scores

Trust Signals

Active community: 1,311 forksCommunity discussions enabled

License: AGPL-3.0

Review license manually.

Commercial use: ✗ Restricted

About

Owner
Palash Debnath (User)
Stars
50,832
Forks
5,634

Explore Further

More tools in the directory

Featured in The Open Source Drop #24