Blog | Documentation | Quick Start | Cookbook | SGLang | Join Slack
⭐ Star SGLang-Omni to help more builders discover open infrastructure for multimodal and speech serving!
- [2026/08] 🎵 Day-0 support for MiniMax Music 3: lyrics + caption → 32 kHz stereo song on
/v1/audio/speech. [Cookbook] - [2026/08] 🚀 SGLang-Omni v0.1.1 is on PyPI. Install with
uv pip install "sglang-omni==0.1.1". [Installation] [Release notes] - [2026/08] 🚀 TTS architecture refactor: shared pipeline state, engine construction, reference encoding, capability metadata, and vocoder scheduling. [Roadmap] [Blog]
- [2026/06] 🔥 MOSS-TTS Local Transformer v1.5 on SGLang-Omni with native-streaming 48 kHz speech. [Blog] [Cookbook]
- [2026/06] 🔥 Higgs Audio v3 TTS for real-time, controllable speech. [Blog] [Cookbook]
SGLang-Omni is a multi-stage serving runtime for omni, speech, and TTS models. Its design target is multi-stage decoding: generation split across heterogeneous stages with different compute patterns, dependency structures, and resource needs. SGLang-Omni owns the pipeline topology, stage lifecycle, inter-stage transport, model-family integration layer, and OpenAI-compatible serving surface, while composing with SGLang for high-performance autoregressive scheduling and model execution where applicable.
- Multi-stage runtime: SGLang-Omni models generation as coordinated stages: preprocessing, encoders, autoregressive engines, talkers, decoders, vocoders, and aggregators.
- Stage-specialized scheduling: Each stage runs behind a scheduler matched to its workload, from SGLang-backed autoregressive scheduling to lightweight preprocessing and streaming vocoder loops.
- Transport-aware execution: A control plane coordinates requests while the relay data plane moves tensor payloads across shared-memory, NCCL, NIXL, and Mooncake backends.
- API surface: OpenAI-compatible endpoints expose multimodal chat, speech generation, batch speech, streaming speech, uploaded voices, and transcription.
- Omni chat and speech: Qwen3-Omni, Ming-Omni — multimodal in, text/audio out.
- Music generation: MiniMax Music 3 — lyrics + caption → 32 kHz stereo song.
- Speech generation: Higgs Audio v3, MOSS-TTS, MOSS-TTS Local, Fish Speech S2-Pro, Qwen3-TTS, Voxtral TTS, Ming-Omni-TTS, dots.tts, ZONOS2 —
/v1/audio/speech, batch, streaming, uploaded voices. - Audio transcription and diarization: Qwen3-ASR, Fun-ASR, ARK-ASR, MOSS-Transcribe-Diarize via
/v1/audio/transcriptions. MOSS-TD supports speaker labels and timestamps (response_format=verbose_json). - SGLang-Omni Router: Multi-worker OpenAI-compatible front door — health, readiness, lifecycle, capability discovery. Router guide.
Additional model guides, including experimental and research-oriented paths, are available in the Cookbook.
- Installation
- TTS usage
- Qwen3-Omni usage
- Qwen3-ASR cookbook
- MOSS-Transcribe-Diarize cookbook
- Omni router
- Developer reference
SGLang-Omni welcomes contributors working on inference systems, kernels, scheduling, inter-stage communication, model runners and cache efficiency, model integration, benchmarking, production deployment. Join the SGLang Slack or read the developer reference.
Organizations interested in supporting SGLang-Omni, TTS, or omni model serving can contact Chenyang Zhao at zhaochenyang@lmsys.org.
SGLang-Omni builds on the SGLang ecosystem and on open model work from the TTS, speech, and omni-model communities. We thank the model teams, systems contributors, and partner organizations helping make open multimodal serving faster, more reliable, and easier to extend.