The most powerful local music generation model that outperforms almost all commercial alternatives, supporting Mac, AMD, Intel, and CUDA devices.
🎙️ 语音与音频
语音识别、语音合成、音色克隆与音乐生成
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingu
A PyTorch-based Speech Toolkit
Very low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces
实时语音对话模型,全双工低延迟。
神经网络说话人分离工具包。
A TTS that fits in your CPU (and pocket)
多语言语音识别、情感识别与事件检测一体模型。
仅 82M 参数的开源 TTS,质量出众。
Transcribe on your own!
An open-source model family for long-form speech, dialogue synthesis, voice design, sound effects, and real-time streaming TTS
Generative models for conditional audio generation
A single Gradio + React WebUI with extensions for ACE-Step, OmniVoice, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS, Stable Audio, MMS, StyleTTS2, MAGNet, AudioGen, MusicGen, To
ggml speech-to-text inference for 16+ model families
Magenta RealTime 2: An Open-Weights Live Music Model
NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.
AuK: An Open-Source Foundational Model for Speech Generation and Editing
Whisper command line client compatible with original OpenAI client based on CTranslate2.
AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.
Jarvis AI Assistant - Voice-powered AI assistant for Mac