VoiceStudio:本地运行的 ElevenLabs 替代品,16 个 TTS 引擎,646 种语言

VoiceStudio: Local-First ElevenLabs Alternative — 16 TTS Engines, 646 Languages

Tech-News #TTS#ASR#voice cloning#local AI#open source#ElevenLabs#audio#speech#MCP#dubbing#audiobook
更新于
🇨🇳 中文

本地运行的语音工作台

ElevenLabs 的核心能力——声音克隆、视频配音、Audiobook 生成——你现在可以在自己的机器上跑,不需要账号、API Key、订阅,也没有用量计费。

VoiceStudio(曾用名 OmniVoice-Studio)是一个开源、全本地的语音工作台,把 16 个 TTS 引擎和 11 个 ASR 引擎统一到一个桌面应用里,支持 646 种语言,覆盖语音 AI 的六个核心工作流。

⭐ 12,712,fork 1,969,AGPL-3.0,Python + Bun/Tauri 构建,持续更新到 2026 年 8 月。


六个工作流

工作流能做什么
声音克隆3-15 秒参考片段,零样本克隆目标声音
声音设计从年龄、口音、音调、风格、表达方式描述生成全新声音
视频配音转录 → 翻译 → 保持说话人 → 合成 → 导出视频
转录 / 听写系统全局快捷键,实时转录,可选本地 LLM 润色
故事与 Audiobook多声音脚本、EPUB/PDF 导入、章节渲染、.m4b 导出
批量队列大规模音频和视频任务并行处理,逐任务进度追踪

核心数据路径是本地的——音频和文字不经过任何第三方服务器。联网功能(远程 worker、模型下载)是明确的可选项,不是默认行为。


引擎生态:16 TTS + 11 ASR

VoiceStudio 的竞争优势不是一个单一模型,而是把当前最好的开源语音模型统一到一个界面里,按需安装、随时切换(Ctrl/Cmd+E)。

TTS 引擎(16 个)

引擎语言数声音克隆macOS ARM许可
VoiceStudio(默认,基于 k2-fsa/OmniVoice)600+✅MPSAGPL-3.0 / Apache-2.0
OmniVoice GGUF600+✅MPS/CPUAGPL-3.0 / Apache-2.0
CosyVoice 39+18方言✅CPUApache-2.0
GPT-SoVITS5✅—MIT
VoxCPM230✅MPSApache-2.0
IndexTTS 2.5ZH/EN/JA/ES/AR✅CPUBilibili 模型许可
MLX-Audio模型相关部分MLX各异
MOSS-TTS-Nano20✅CPUApache-2.0
Sherpa-ONNX20+—CPUApache-2.0
KittenTTS英语—CPUMIT
PocketTTS6种欧洲语言✅CPUCC-BY-4.0(需授权)
Supertonic 331—CPUOpenRAIL-M
MOSS-TTS-v1.531✅CPUApache-2.0
dots.tts24✅CPUApache-2.0
Confucius4-TTS14✅CPUApache-2.0
MOSS-TTS-v1.531✅CPUApache-2.0

默认引擎 VoiceStudio(OmniVoice)支持 600+ 语言、声音克隆、指令式合成,是开箱即用的最全能选项。没有克隆能力的引擎在视频配音和固定声音批量任务里会被直接拒绝(而不是悄悄换引擎),确保输出的可预期性。

ASR 引擎(11 个)

涵盖 Whisper 系列、WhisperX、Pyannote 说话人分离、实时流式识别等,配合 TTS 引擎构成完整的语音处理管线。


技术架构

层技术
前端 / 桌面Tauri + Bun(TypeScript)
后端Python(uv 管理依赖)
计算CUDA · Apple Silicon MPS/MLX · ROCm · CPU
接口本地 REST/SSE/WebSocket API · OpenAI 兼容音频 API · MCP Server
模型管理内置 Model Catalogue,在线安装/卸载/路由,支持远程 Worker

MCP Server 是一个值得关注的细节:VoiceStudio 暴露合成和转录工具给任何 MCP 客户端(Claude、Cursor 等),意味着你可以在 AI 编码工具里直接调用本地语音合成——不用离开工作区。


对比 ElevenLabs

VoiceStudioElevenLabs
数据路径本地(音频和文字不出机器)经过 ElevenLabs 服务器
费用免费(你提供算力)订阅 / 按用量计费
离线使用✅(模型下载后)❌
引擎选择16 个 TTS + 11 个 ASR闭源,固定
语言支持646 种(取决于引擎)32 种
定制性开源、可改引擎、可改路由有限
维护你自己管更新和算力供应商管基础设施

适合 VoiceStudio 的场景:私有数据(法律/医疗/企业内容)、离线环境、高频批量生产(不想按量付费)、自研工作流集成。


硬件需求

最低推荐
OSWindows 10 x64 · macOS 13.3+ Apple Silicon · Linux x86_64当前系统版本
RAM8 GB16 GB+
磁盘10 GB20 GB+ SSD
GPU可选(支持纯 CPU 模式)NVIDIA CUDA 或 Apple Silicon
VRAM4 GB(使用 GPU 时)8 GB+(大型引擎更多)

注意:Intel Mac 无法运行本地 Python 后端,只能连接远程 Worker。Apple Silicon 是 macOS 上的原生平台。


安装与快速开始

从 GitHub Releases 下载对应平台的安装包(DMG / MSI / AppImage),首次启动自动创建 Python 环境并下载默认模型,后续启动复用缓存。

首次声音克隆三步:

  1. 打开 VoiceStudio → Voice Cloning
  2. 上传干净的参考音频(3 秒可用,5-15 秒效果更好)
  3. 输入文字,选语言,点 Generate

从源码运行:

git clone https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio
bun install
bun run desktop       # 桌面端
# bun run dev         # 浏览器 UI

总结

VoiceStudio 做了一件看起来简单但执行门槛很高的事:把语音 AI 的主流开源模型(16 个 TTS、11 个 ASR)统一到一个本地桌面应用里,覆盖从声音克隆到 Audiobook 生产的完整工作流,并暴露 OpenAI 兼容 API 和 MCP Server 给工具链集成。

12k Star,1.9k Fork,ElevenLabs 的替代品——不是功能对等,而是本地优先、数据自控、不按量计费。

GitHub: debpalash/VoiceStudio ⭐12712
官网: voicestudio.sh
Discord: discord.gg/bzQavDfVV9
许可: AGPL-3.0

🇬🇧 English

VoiceStudio: Local-First ElevenLabs Alternative

ElevenLabs’ core capabilities — voice cloning, video dubbing, audiobook generation — can now run on your own machine, with no account, API key, subscription, or usage meter.

VoiceStudio (formerly OmniVoice-Studio) is an open-source, fully-local voice workstation that unifies 16 TTS engines and 11 ASR engines in a single desktop app, supports 646 languages, and covers six core voice AI workflows.

⭐12,712, 1,969 forks, AGPL-3.0, Python + Bun/Tauri, actively updated through August 2026.


Six Workflows

WorkflowWhat it does
Voice CloningZero-shot clone from a 3–15 second reference clip
Voice DesignGenerate a new voice from age, accent, pitch, style, and delivery instructions
Video DubbingTranscribe → translate → preserve speakers → synthesize → export video
DictationSystem-wide shortcut, live transcription, optional local LLM cleanup
Stories & AudiobooksMulti-voice scripts, EPUB/PDF import, chapter rendering, .m4b export
Batch QueueLarge-scale audio and video job processing with per-job progress tracking

The core data path is local — audio and text never reach a third-party server. Network-backed features (remote workers, model downloads) are explicit opt-ins, not defaults.


Engine Ecosystem: 16 TTS + 11 ASR

VoiceStudio’s competitive advantage isn’t a single model — it’s unifying the best open-source voice models into one interface with on-demand installation and instant switching (Ctrl/Cmd+E).

TTS Engines (16)

The default engine — VoiceStudio (powered by k2-fsa/OmniVoice) — supports 600+ languages, voice cloning, and instruction-driven synthesis. Engines without cloning support are rejected rather than silently swapped in dubbing and pinned-voice batch jobs, keeping outputs predictable.

Key highlights:

  • OmniVoice / OmniVoice GGUF — 600+ languages, clone + instruct, Apple Silicon MPS support
  • CosyVoice 3 — 9 languages + 18 Chinese dialects, clone + instruct
  • GPT-SoVITS — MIT, 5 languages, popular for high-quality clone
  • IndexTTS 2.5 — Chinese/English/Japanese/Spanish/Arabic
  • MLX-Audio — MLX native on Apple Silicon
  • Sherpa-ONNX — lightweight, 20+ languages, CPU-first

ASR Engines (11)

Covers Whisper variants, WhisperX with speaker diarization (Pyannote), real-time streaming recognition, and more — a full audio processing pipeline alongside TTS.


Technical Architecture

LayerTechnology
Frontend / DesktopTauri + Bun (TypeScript)
BackendPython (uv-managed)
ComputeCUDA · Apple Silicon MPS/MLX · ROCm · CPU
InterfacesLocal REST/SSE/WebSocket API · OpenAI-compatible audio API · MCP Server
Model managementBuilt-in Model Catalogue: install/remove/route, remote worker support

MCP Server is worth highlighting: VoiceStudio exposes synthesis and transcription tools to any MCP client (Claude, Cursor, etc.), letting you call local voice synthesis directly from within AI coding tools — without leaving your workspace.


Compared to ElevenLabs

VoiceStudioElevenLabs
Data pathLocal by defaultProcessed by ElevenLabs servers
CostFree (you supply compute)Subscription or metered API
Offline✅ after model download❌
Engine choice16 TTS + 11 ASRClosed-source, fixed
Language support646 (engine-dependent)32
CustomizationOpen source, swap engines, modify routingProvider-limited
MaintenanceYou manage updates and computeProvider manages infrastructure

Best fit for VoiceStudio: private data (legal, medical, enterprise), offline environments, high-volume batch production (no per-character billing), and custom toolchain integration.


Hardware Requirements

MinimumRecommended
OSWindows 10 x64 · macOS 13.3+ Apple Silicon · Linux x86_64Current supported OS release
RAM8 GB16 GB+
Disk10 GB20 GB+ SSD
GPUOptional (CPU mode supported)NVIDIA CUDA or Apple Silicon
VRAM4 GB if using GPU8 GB+ (large engines need more)

Note: Intel Macs cannot run the local Python backend — connect a remote worker instead. Apple Silicon is the native macOS platform.


Install and Quick Start

Download the installer for your platform from GitHub Releases (DMG / MSI / AppImage). First launch creates a managed Python environment and downloads the default model; subsequent launches reuse the cache.

First voice clone in three steps:

  1. Open VoiceStudio → Voice Cloning
  2. Add a clean reference clip (3 seconds works; 5–15 seconds usually better)
  3. Enter text, choose a language, click Generate

Run from source:

git clone https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio
bun install
bun run desktop    # desktop app
# bun run dev      # browser UI

Summary

VoiceStudio does something that sounds simple but has a high execution bar: it unifies the major open-source voice models (16 TTS, 11 ASR) into a local desktop app covering the complete workflow from voice cloning to audiobook production, then exposes an OpenAI-compatible API and MCP Server for toolchain integration.

12k stars, 1.9k forks. An ElevenLabs alternative — not feature-parity in every edge case, but local-first, data-controlled, and no per-character billing.

GitHub: debpalash/VoiceStudio ⭐12712
Website: voicestudio.sh
Discord: discord.gg/bzQavDfVV9
License: AGPL-3.0

💬 评论与讨论

使用 GitHub 账号登录后发表评论

关于本站 · 免责声明

🍄 Mushroom Research Blog 是非营利、免费公开的个人科技观察博客与公众号 XStack18,不接受商业合作、不代表任何企业或机构立场,也不谋求商业利益。我们以个人视角客观中立地记录和分析 AI、Web3 等领域的最新模型发布与技术动态——不止转述新闻标题或二手信息,而是给出有独立思考的深入分析,希望帮更多人获得有价值的一手科技认知。

⚠️ 文中介绍的开源代码与模型,仅供学习交流与技术借鉴。它们大多仍处于早期阶段,有待进一步研究和验证,请勿直接用于工作或生产环境;如需采用,请先自行充分测试,并核实其许可证与安全性。
Open-source code and models featured here are shared for learning and reference only. Most are early-stage and still need further study and verification — please don't use them directly in your work or in production. Test them thoroughly and check their licenses and security first.

  1. 本站文章均为作者基于公开信息的个人研究与观点整理,不代表文中提及的任何公司、产品、模型的官方立场,未与其构成商业关联或合作关系。
  2. 科技行业信息更新极快,我们尽力保证内容准确、及时,但不对完整性、实时性做绝对保证,具体请以相关企业/项目官方公告为准。
  3. 文中引用的第三方商标、产品名称、图片、数据等版权归原权利人所有,我们会尽量注明来源;如你认为存在版权疑问或侵权,请通过下方邮箱联系我们,收到通知后会尽快核实处理(更正、加注来源或删除)。
  4. 文章内容仅为技术科普与个人观点,不构成投资、法律或其他专业建议,据此进行任何决策的后果需自行判断和承担。

📮 侵权 / 勘误 / 合作咨询:hello@mushroom.cv