⚡ Connect multiple TTS backends with a unified interface — Minimalist text-to-speech orchestration supporting offline and cloud providers.
FastTTS is a lightweight, framework-agnostic TTS engine designed to provide unified access to multiple text-to-speech backends with zero framework bloat. It supports Piper (offline), Windows SAPI (system), ElevenLabs (cloud), and Deepgram (cloud) through a single clean API.
Watch Demo (YouTube) | Watch JMH Benchmark (YouTube)
import fasttts.FastTTS;
import fasttts.backends.piper.PiperBackend;
import fasttts.backends.windows.WindowsTTSBackend;
import fasttts.backends.elevenlabs.ElevenLabsBackend;
import fasttts.backends.deepgram.DeepgramBackend;
import fasttts.core.FastTTSAudio;
public class Demo {
public static void main(String[] args) throws Exception {
FastTTS tts = new FastTTS();
// Windows SAPI (no setup required)
tts.registerBackend(new WindowsTTSBackend());
FastTTSAudio audio = tts.speak("Hello World");
// Piper (offline, requires piper.exe and models)
tts.registerBackend(new PiperBackend("piper.exe", "models/de_DE-thorsten-medium.onnx"));
FastTTSAudio germanAudio = tts.speak("Hallo Welt");
// ElevenLabs (cloud, requires API key)
tts.registerBackend(new ElevenLabsBackend("your-api-key"));
FastTTSAudio cloudAudio = tts.speak("Hello World");
// Deepgram (cloud, requires API key)
tts.registerBackend(new DeepgramBackend("your-api-key"));
FastTTSAudio deepgramAudio = tts.speak("Hello World");
}
}- Why FastTTS?
- Key Features
- Real-World Use Cases
- Performance Benchmarks
- Architecture Overview
- API Quick Reference
- Technical Demos & Benchmarks
- Installation
- Backend Setup
- Documentation
- Platform Support
- License
- Related Projects
Traditional TTS libraries force developers into heavyweight Python dependencies, complex cloud API integrations, or platform-specific code. FastTTS provides:
- 100% Native JVM Pipeline — Orchestrates multiple TTS backends in a single JVM process with unified interface.
- Offline and Cloud Support — Seamlessly switch between local models (Piper) and cloud providers (ElevenLabs, Deepgram).
- Model Agnostic — Works with offline ONNX models, system voices, and cloud APIs through the same
FastTTSBackendinterface. - Zero Configuration Overlap — Integrates seamlessly with existing FastJava ecosystem libraries.
FastTTS unifies local neural inference, system voices, and cloud TTS into a single framework-agnostic API:
| Feature | MaryTTS (Legacy Java) | Python Coqui/Piper Subprocess | FastTTS |
|---|---|---|---|
| Synthesis Backends | Obsolete Java voices only | Single CLI wrapper | Unified (SAPI, Piper, ElevenLabs, Deepgram) |
| Offline Privacy | Yes (Robotic sound) | Yes (Heavy Python runtime) | 100% Offline (Piper ONNX / Windows SAPI) |
| Startup / Synthesis | 800–2500 ms (Heap heavy) | 2000–5000 ms (Process spawn) | ~255 ms (Windows SAPI) / Fast Piper |
| Voice Quality | Synthetic 2000s robotic | Neural quality | State-of-the-Art Neural & System |
| Barge-In Ready | Difficult to interrupt | Stalled CLI process | Instant Cancellation via FastVAD |
| Dependencies | Massive legacy JARs | Python 3 + Pip dependencies | Pure Java 17+ backed by FastCore |
- 🎭 Multiple Backend Support — Unified interface for Piper (offline), Windows SAPI (system), ElevenLabs (cloud), and Deepgram (cloud).
- 📱 Offline Capable — Run TTS locally with Piper models without internet connection.
- ☁️ Cloud Integration — Access high-quality cloud voices from ElevenLabs and Deepgram.
- ⚡ Performance Focused — Built for low-latency synthesis with detailed timing metrics.
- 🔌 Simple API — Clean, intuitive interface for text-to-speech synthesis.
- 🗣️ Conversational Voice AI Assistants: Low-latency neural speech synthesis for AI chatbots and desktop voice assistants.
- 📖 Audiobook & Content Reader: High-speed offline speech synthesis for document reading and accessibility tools.
- 📢 In-App & Game Audio Notifications: Real-time voice announcements with zero Garbage Collection latency impact.
- 🌐 Multi-Language Accessibility Engines: Seamlessly switch between local Piper voices and cloud APIs (ElevenLabs, Deepgram).
FastTTS is built for high-performance text-to-speech synthesis. Based on the built-in demo timing metrics:
Backend Load Time Synth Time Total Time
Windows SAPI 104 ms 151 ms 255 ms
Piper (offline) 0 ms 1761 ms 1761 ms
Windows SAPI provides the fastest synthesis (255ms total) for quick feedback, while Piper offers higher quality offline synthesis (1761ms total) with no internet dependency.
Windows SAPI Backend
System-native text-to-speech using Windows Speech API (SAPI). No external dependencies required.
Piper Backend
Offline TTS using the Piper neural TTS engine. Requires piper.exe and ONNX model files for local synthesis.
ElevenLabs Backend
Cloud-based TTS accessing ElevenLabs high-quality voices. Requires API key and internet connection.
Deepgram Backend
Cloud-based TTS using Deepgram's fast synthesis API. Requires API key and internet connection.
FastTTS (This Library — The Orchestration Layer)
Higher-level TTS framework that provides a unified interface for all backends, allowing seamless switching between offline and cloud providers.
| Method / Signature | Return Type | Description | Docs |
|---|---|---|---|
registerBackend(FastTTSBackend backend) |
void |
Registers a TTS backend with the orchestrator. | Wiki |
speak(String text) |
FastTTSAudio |
Synthesizes text to audio using the active backend. | Wiki |
speak(String backend, String text, FastTTSVoice voice, FastTTSConfig config) |
FastTTSAudio |
Synthesizes text with specific backend and configuration. | Wiki |
stream(String backend, String text, FastTTSVoice voice, FastTTSConfig config, Consumer<byte[]> consumer) |
void |
Streams audio chunks directly to a byte consumer callback. | Wiki |
use(String backendName) |
void |
Sets the active backend by name. | Wiki |
getAllVoices() |
List<FastTTSVoice> |
Returns available voices across all registered backends. | Wiki |
getBackend(String name) |
FastTTSBackend |
Retrieves a registered backend by name. | Wiki |
| Case | Java Example | Launcher | Description |
|---|---|---|---|
| Multi-Backend CLI Demo | Demo.java | run-demo.bat |
Command-line speech synthesis across Windows SAPI, Piper, ElevenLabs, and Deepgram. |
| JMH Microbenchmark Suite | Benchmark.java | run-benchmark.bat |
Formal OpenJDK JMH throughput and latency benchmarks for TTS synthesis. |
Add the JitPack repository and the complete dependency stack to your pom.xml:
<repositories>
<repository>
<id>jitpack.io</id>
<url>https://jitpack.io</url>
</repository>
</repositories>
<dependencies>
<!-- FastTTS Engine -->
<dependency>
<groupId>com.github.andrestubbe</groupId>
<artifactId>FastTTS</artifactId>
<version>0.1.2</version>
</dependency>
<!-- FastSIMD Hardware Vector Acceleration Engine -->
<dependency>
<groupId>com.github.andrestubbe</groupId>
<artifactId>FastSIMD</artifactId>
<version>0.1.3</version>
</dependency>
<!-- FastMemory Aligned Allocator -->
<dependency>
<groupId>com.github.andrestubbe</groupId>
<artifactId>FastMemory</artifactId>
<version>0.1.1</version>
</dependency>
<!-- FastPointer Address Wrapper -->
<dependency>
<groupId>com.github.andrestubbe</groupId>
<artifactId>FastPointer</artifactId>
<version>0.1.1</version>
</dependency>
<!-- FastAudioProcess Audio Engine -->
<dependency>
<groupId>com.github.andrestubbe</groupId>
<artifactId>FastAudioProcess</artifactId>
<version>0.1.1</version>
</dependency>
<!-- FastCore Unified JNI Loader -->
<dependency>
<groupId>com.github.andrestubbe</groupId>
<artifactId>FastCore</artifactId>
<version>0.1.0</version>
</dependency>
</dependencies>repositories {
maven { url 'https://jitpack.io' }
}
dependencies {
implementation 'com.github.andrestubbe:FastTTS:0.1.2'
implementation 'com.github.andrestubbe:FastSIMD:0.1.3'
implementation 'com.github.andrestubbe:FastMemory:0.1.1'
implementation 'com.github.andrestubbe:FastPointer:0.1.1'
implementation 'com.github.andrestubbe:FastAudioProcess:0.1.1'
implementation 'com.github.andrestubbe:FastCore:0.1.0'
}- No installation required — uses Windows built-in voices
- Works immediately after FastTTS installation
- Multiple voices available (system default)
- Download Piper: https://github.com/rhasspy/piper/releases
- Extract Piper to a directory (e.g.,
C:\Piper\) - Set environment variable:
set PIPER_PATH=C:\Piper\piper.exe - Or place
piper.exein your project directory - Download voice models (
.onnxand.onnx.json) from: https://huggingface.co/models?search=piper - Place model files in the
models/folder
Supported Piper ONNX Voices (Examples):
de_DE-thorsten-medium.onnx— High-quality German male voice (bundled inmodels/)en_US-lessac-medium.onnx— Clear US English voiceen_US-amy-medium.onnx— Expressive US English voice
- Requires API key from: https://elevenlabs.io
- High-quality voices
- Cloud-based (requires internet)
- Requires API key from: https://deepgram.com
- Fast cloud-based synthesis
- Multiple voice options
- CHANGELOG.md: Release notes and version history.
- REFERENCE.md: Core API reference manual.
- PHILOSOPHY.md: Engineering rationale for zero-allocation performance.
- COMPILE.md: Full compilation guide (MSVC C++17 build chain + JNI Setup).
- ROADMAP.md: Future development goals.
| Platform | Architecture | Status | Notes |
|---|---|---|---|
| Windows 10/11 | x64 | ✅ Fully Supported | Native Windows SAPI and MSVC AVX2 compilation |
| Linux | x64, ARM64 | 🚧 Planned | Native Piper and ALSA / Pulse audio support |
| macOS | Apple Silicon, x64 | 🚧 Planned | AVFoundation / Piper support |
MIT License — See LICENSE file for details.
- FastCore — Native JNI loader for FastJava libraries
- FastAudioPlayer — Native audio playback for Java via WASAPI
- FastAI — Unified lightweight AI model client interface
- FastAIModel — Embedded GGUF and ONNX runtimes for local feature embeddings
- FastAIRag — Unified, zero-bloat RAG pipeline client for Java
- FastAIVectorDB — High-speed native C++ SIMD vector database
- FastContentParse — Standardized Java document parser for text extraction
- FastContentChunk — High-performance native SIMD tokenizer and chunker
Part of the FastJava Ecosystem — Making the JVM faster. Small package. Maximum speed. Zero bloat. 🚀📋
