Manage your Prompts with PROMPT01 Use "THEJOAI" Code 50% OFF

Mid Takes Only

Mid Takes Only
Launch Date: Aug. 19, 2026
Pricing: No Info
TTS API, Voice Cloning, Conversational AI, Open Source Models, Real-Time Speech

Real-Time Speech for Voice Agents: The nineninesix.ai Solution

Overview

nineninesix.ai provides sub-100ms streaming text-to-speech (TTS) designed specifically for voice agents. The service offers human-like audio quality across five languages, delivering approximately 22 hours of audio for $5. Access is available via a single API call, supporting both cloud-based API usage and open-weight model downloads.

Benefits

The platform distinguishes itself through exceptional latency, measured in two ways: vendor-reported model-only latency and independent end-to-end latency.

Model-Only Latency (Vendor Claims)

This metric excludes network overhead and represents the time taken by the model to generate audio once text is received.*nineninesix Gepard 1.0:~35 ms*ElevenLabs Flash v2.5:~75 ms*Cartesia Sonic 3.5:<90 ms*Deepgram Aura-2:~90 ms*OpenAI gpt-4o-mini-tts:~230 ms*Alibaba Qwen3 TTS Flash:~300 ms*xAI Grok TTS:No model-only figure published.

End-to-End Latency (Independent Coval P50 Benchmark)

This metric includes network latency and represents the total time from request to first audio byte (TTFA).*nineninesix Gepard 1.0:90 ms* (Self-measured, pending independent verification)*ElevenLabs Flash v2.5:207 ms*Cartesia Sonic 3.5:288 ms*Deepgram Aura-2:339 ms*xAI Grok TTS:393 ms*Alibaba Qwen3 TTS Flash:698 ms*OpenAI gpt-4o-mini-tts:800 ms

Note: Latency measurements vary by vendor methodology. These figures should be treated as directional indicators rather than exact specifications.

Product Features

Gepard 1.0 Model

nineninesix has open-sourcedGepard 1.0, a 555M parameter streaming TTS model built on Qwen 3.5. Key characteristics include:*Streaming Capability:Starts generating audio as text arrives (autoregressive generation).*Speed:Approximately 50ms to Time To First Audio (TTFA) and 20x Real-Time Factor (RTF) on a single RTX 5090.*Voice Cloning:Capable of cloning voices from just a few seconds of audio.*Efficiency:Uses Safetensors for efficient inference and is native to vLLM.*License:Distributed under the Apache 2.0 license.*Availability:Weights and demos are available on Hugging Face.

Additional Models

The platform also highlightskani-tts-370m, a lightweight open-source model featuring:*Size:370M parameters, optimized for consumer GPUs.*Architecture:Utilizes NanoCodec and LFM2-350M.*Quality:Trained with modern neural TTS techniques for natural and expressive voices.*Integration:Includes a ComfyUI node for Windows users.

Pricing and Limits

Billing is based on a simple postpaid model where 1 credit equals 1 character of input text (including emojis and accented glyphs). Charges apply only after successful generation; failed generations are never billed.

Rate Limits (Per Organization)

Limits are shared across all API keys within an organization. Extra keys are intended for rotation and security, not for increasing capacity.

PlanRequests / MinConcurrent StreamsCost
Tier 1 (Default)605Free
Tier 260025$50 lifetime
Tier 3 (High Volume)6,000100Custom

Concurrency Mechanics

Concurrency is measured per active-speech turn, not per API call. A stream counts against the limit only while actively generating audio (roughly 0.5–4 seconds per turn). Idle time (user speaking, silence, or thinking) costs nothing. This architecture allows a single concurrency slot to serve multiple live calls simultaneously.

API and Deployment

Users can interact with the models via:*Cloud API:Simple HTTP calls (/tts/bytes,/tts/sse) or long-lived WebSocket connections (/tts/websocket).*Open Weights:Download models directly from Hugging Face for self-hosting.

The service is designed to be "Agent Ready," facilitating seamless integration into conversational AI workflows where low latency is critical to maintaining natural dialogue flow.

NOTE:

This content is either user submitted or generated using AI technology (including, but not limited to, Google Gemini API, Llama, Grok, and Mistral), based on automated research and analysis of public data sources from search engines like DuckDuckGo, Google Search, and SearXNG, and directly from the tool's own website and with minimal to no human editing/review. THEJO AI is not affiliated with or endorsed by the AI tools or services mentioned. This is provided for informational and reference purposes only, is not an endorsement or official advice, and may contain inaccuracies or biases. Please verify details with original sources.

Comments

Loading...