Skip to content
andrestubbePublic

About

🔊 High‑performance native Text‑to‑Speech engine for Java — ultra‑low latency synthesis via Windows SAPI/WinRT, Piper, and cloud backends like Eevenlabs, Deepgram or OpenAI with real‑time streaming support.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

FastTTS 0.1.2 [ALPHA-2026-09-30] — Unified, Zero-Bloat TTS Backend Orchestration for Java

Status License: MIT Java Platform JitPack


⚡ Connect multiple TTS backends with a unified interface — Minimalist text-to-speech orchestration supporting offline and cloud providers.

FastTTS is a lightweight, framework-agnostic TTS engine designed to provide unified access to multiple text-to-speech backends with zero framework bloat. It supports Piper (offline), Windows SAPI (system), ElevenLabs (cloud), and Deepgram (cloud) through a single clean API.

Watch Demo (YouTube) | Watch JMH Benchmark (YouTube)

FastTTS Showcase


Quick Start — Example

import fasttts.FastTTS;
import fasttts.backends.piper.PiperBackend;
import fasttts.backends.windows.WindowsTTSBackend;
import fasttts.backends.elevenlabs.ElevenLabsBackend;
import fasttts.backends.deepgram.DeepgramBackend;
import fasttts.core.FastTTSAudio;

public class Demo {
    public static void main(String[] args) throws Exception {
        FastTTS tts = new FastTTS();
        
        // Windows SAPI (no setup required)
        tts.registerBackend(new WindowsTTSBackend());
        FastTTSAudio audio = tts.speak("Hello World");
        
        // Piper (offline, requires piper.exe and models)
        tts.registerBackend(new PiperBackend("piper.exe", "models/de_DE-thorsten-medium.onnx"));
        FastTTSAudio germanAudio = tts.speak("Hallo Welt");
        
        // ElevenLabs (cloud, requires API key)
        tts.registerBackend(new ElevenLabsBackend("your-api-key"));
        FastTTSAudio cloudAudio = tts.speak("Hello World");
        
        // Deepgram (cloud, requires API key)
        tts.registerBackend(new DeepgramBackend("your-api-key"));
        FastTTSAudio deepgramAudio = tts.speak("Hello World");
    }
}

Table of Contents


Why FastTTS?

Traditional TTS libraries force developers into heavyweight Python dependencies, complex cloud API integrations, or platform-specific code. FastTTS provides:

  • 100% Native JVM Pipeline — Orchestrates multiple TTS backends in a single JVM process with unified interface.
  • Offline and Cloud Support — Seamlessly switch between local models (Piper) and cloud providers (ElevenLabs, Deepgram).
  • Model Agnostic — Works with offline ONNX models, system voices, and cloud APIs through the same FastTTSBackend interface.
  • Zero Configuration Overlap — Integrates seamlessly with existing FastJava ecosystem libraries.

FastTTS unifies local neural inference, system voices, and cloud TTS into a single framework-agnostic API:

Feature MaryTTS (Legacy Java) Python Coqui/Piper Subprocess FastTTS
Synthesis Backends Obsolete Java voices only Single CLI wrapper Unified (SAPI, Piper, ElevenLabs, Deepgram)
Offline Privacy Yes (Robotic sound) Yes (Heavy Python runtime) 100% Offline (Piper ONNX / Windows SAPI)
Startup / Synthesis 800–2500 ms (Heap heavy) 2000–5000 ms (Process spawn) ~255 ms (Windows SAPI) / Fast Piper
Voice Quality Synthetic 2000s robotic Neural quality State-of-the-Art Neural & System
Barge-In Ready Difficult to interrupt Stalled CLI process Instant Cancellation via FastVAD
Dependencies Massive legacy JARs Python 3 + Pip dependencies Pure Java 17+ backed by FastCore

Key Features

  • 🎭 Multiple Backend Support — Unified interface for Piper (offline), Windows SAPI (system), ElevenLabs (cloud), and Deepgram (cloud).
  • 📱 Offline Capable — Run TTS locally with Piper models without internet connection.
  • ☁️ Cloud Integration — Access high-quality cloud voices from ElevenLabs and Deepgram.
  • ⚡ Performance Focused — Built for low-latency synthesis with detailed timing metrics.
  • 🔌 Simple API — Clean, intuitive interface for text-to-speech synthesis.

Real-World Use Cases

  • 🗣️ Conversational Voice AI Assistants: Low-latency neural speech synthesis for AI chatbots and desktop voice assistants.
  • 📖 Audiobook & Content Reader: High-speed offline speech synthesis for document reading and accessibility tools.
  • 📢 In-App & Game Audio Notifications: Real-time voice announcements with zero Garbage Collection latency impact.
  • 🌐 Multi-Language Accessibility Engines: Seamlessly switch between local Piper voices and cloud APIs (ElevenLabs, Deepgram).

Performance Benchmarks

FastTTS is built for high-performance text-to-speech synthesis. Based on the built-in demo timing metrics:

Backend          Load Time    Synth Time    Total Time
Windows SAPI     104 ms       151 ms        255 ms
Piper (offline)  0 ms         1761 ms       1761 ms

Windows SAPI provides the fastest synthesis (255ms total) for quick feedback, while Piper offers higher quality offline synthesis (1761ms total) with no internet dependency.


Architecture Overview

Windows SAPI Backend
System-native text-to-speech using Windows Speech API (SAPI). No external dependencies required.

Piper Backend
Offline TTS using the Piper neural TTS engine. Requires piper.exe and ONNX model files for local synthesis.

ElevenLabs Backend
Cloud-based TTS accessing ElevenLabs high-quality voices. Requires API key and internet connection.

Deepgram Backend
Cloud-based TTS using Deepgram's fast synthesis API. Requires API key and internet connection.

FastTTS (This Library — The Orchestration Layer)
Higher-level TTS framework that provides a unified interface for all backends, allowing seamless switching between offline and cloud providers.


API Quick Reference

Method / Signature Return Type Description Docs
registerBackend(FastTTSBackend backend) void Registers a TTS backend with the orchestrator. Wiki
speak(String text) FastTTSAudio Synthesizes text to audio using the active backend. Wiki
speak(String backend, String text, FastTTSVoice voice, FastTTSConfig config) FastTTSAudio Synthesizes text with specific backend and configuration. Wiki
stream(String backend, String text, FastTTSVoice voice, FastTTSConfig config, Consumer<byte[]> consumer) void Streams audio chunks directly to a byte consumer callback. Wiki
use(String backendName) void Sets the active backend by name. Wiki
getAllVoices() List<FastTTSVoice> Returns available voices across all registered backends. Wiki
getBackend(String name) FastTTSBackend Retrieves a registered backend by name. Wiki

Technical Demos & Benchmarks

Case Java Example Launcher Description
Multi-Backend CLI Demo Demo.java run-demo.bat Command-line speech synthesis across Windows SAPI, Piper, ElevenLabs, and Deepgram.
JMH Microbenchmark Suite Benchmark.java run-benchmark.bat Formal OpenJDK JMH throughput and latency benchmarks for TTS synthesis.

Installation

Option 1: Maven (Recommended)

Add the JitPack repository and the complete dependency stack to your pom.xml:

<repositories>
    <repository>
        <id>jitpack.io</id>
        <url>https://jitpack.io</url>
    </repository>
</repositories>

<dependencies>
    <!-- FastTTS Engine -->
    <dependency>
        <groupId>com.github.andrestubbe</groupId>
        <artifactId>FastTTS</artifactId>
        <version>0.1.2</version>
    </dependency>

    <!-- FastSIMD Hardware Vector Acceleration Engine -->
    <dependency>
        <groupId>com.github.andrestubbe</groupId>
        <artifactId>FastSIMD</artifactId>
        <version>0.1.3</version>
    </dependency>

    <!-- FastMemory Aligned Allocator -->
    <dependency>
        <groupId>com.github.andrestubbe</groupId>
        <artifactId>FastMemory</artifactId>
        <version>0.1.1</version>
    </dependency>

    <!-- FastPointer Address Wrapper -->
    <dependency>
        <groupId>com.github.andrestubbe</groupId>
        <artifactId>FastPointer</artifactId>
        <version>0.1.1</version>
    </dependency>

    <!-- FastAudioProcess Audio Engine -->
    <dependency>
        <groupId>com.github.andrestubbe</groupId>
        <artifactId>FastAudioProcess</artifactId>
        <version>0.1.1</version>
    </dependency>

    <!-- FastCore Unified JNI Loader -->
    <dependency>
        <groupId>com.github.andrestubbe</groupId>
        <artifactId>FastCore</artifactId>
        <version>0.1.0</version>
    </dependency>
</dependencies>

Option 2: Gradle (via JitPack)

repositories {
    maven { url 'https://jitpack.io' }
}

dependencies {
    implementation 'com.github.andrestubbe:FastTTS:0.1.2'
    implementation 'com.github.andrestubbe:FastSIMD:0.1.3'
    implementation 'com.github.andrestubbe:FastMemory:0.1.1'
    implementation 'com.github.andrestubbe:FastPointer:0.1.1'
    implementation 'com.github.andrestubbe:FastAudioProcess:0.1.1'
    implementation 'com.github.andrestubbe:FastCore:0.1.0'
}

Backend Setup

Windows SAPI (System TTS)

  • No installation required — uses Windows built-in voices
  • Works immediately after FastTTS installation
  • Multiple voices available (system default)

Piper (Offline TTS)

Supported Piper ONNX Voices (Examples):

  • de_DE-thorsten-medium.onnx — High-quality German male voice (bundled in models/)
  • en_US-lessac-medium.onnx — Clear US English voice
  • en_US-amy-medium.onnx — Expressive US English voice

ElevenLabs (Cloud TTS)

Deepgram (Cloud TTS)


Documentation


Platform Support

Platform Architecture Status Notes
Windows 10/11 x64 ✅ Fully Supported Native Windows SAPI and MSVC AVX2 compilation
Linux x64, ARM64 🚧 Planned Native Piper and ALSA / Pulse audio support
macOS Apple Silicon, x64 🚧 Planned AVFoundation / Piper support

License

MIT License — See LICENSE file for details.


Related Projects

  • FastCore — Native JNI loader for FastJava libraries
  • FastAudioPlayer — Native audio playback for Java via WASAPI
  • FastAI — Unified lightweight AI model client interface
  • FastAIModel — Embedded GGUF and ONNX runtimes for local feature embeddings
  • FastAIRag — Unified, zero-bloat RAG pipeline client for Java
  • FastAIVectorDB — High-speed native C++ SIMD vector database
  • FastContentParse — Standardized Java document parser for text extraction
  • FastContentChunk — High-performance native SIMD tokenizer and chunker

Part of the FastJava Ecosystem — Making the JVM faster. Small package. Maximum speed. Zero bloat. 🚀📋

About

🔊 High‑performance native Text‑to‑Speech engine for Java — ultra‑low latency synthesis via Windows SAPI/WinRT, Piper, and cloud backends like Eevenlabs, Deepgram or OpenAI with real‑time streaming support.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages