Text-to-Speech (TTS) | Hume API

Octave 2 (preview) and EVI 4-mini are live! Expanded language support and lower latency for faster, more natural responses. Learn more.

Octave TTS is the first text-to-speech system built on LLM intelligence. Octave understands the text it speaks, both emotionally and semantically. It knows when to whisper secrets, when to shout in triumph, and when to calmly state facts. It produces industry-leading voice quality and expressiveness at real-time speeds. Create any voice you can imagine on Octave through prompting, or use Octave to create a state-of-the-art clone of your own voice.

You retain full ownership of any audio content you generate using Octave. For complete details on ownership rights, please see Hume’s Terms of Use.

Features

Key capabilities

Octave versions

Feature Octave 1 Octave 2 (preview)
Supported languages English, Spanish English, Japanese, Korean, Spanish, French, Portuguese, Italian, German, Russian, Hindi, Arabic
Model latency ~200ms ~100ms
Voice cloning
Voice design (English only, multilingual coming soon)
Acting instructions (Coming soon)
Continuation
Timestamps (phoneme/word)

Quickstart

Accelerate your project setup with our comprehensive quickstart guides, designed to integrate Octave TTS into your TypeScript or Python applications. Each guide walks you through API integration and demonstrates text-to-speech synthesis, helping you get up and running quickly.

TypeScript Integrate Octave TTS into web and Node.js applications using our TypeScript SDK.](https://dev.hume.ai/docs/text-to-speech-tts/quickstart/typescript) Python Use our Python SDK to integrate Octave TTS into your Python applications.](https://dev.hume.ai/docs/text-to-speech-tts/quickstart/python) .NET Use our .NET SDK to integrate Octave TTS into your .NET applications.](https://dev.hume.ai/docs/text-to-speech-tts/quickstart/dotnet) CLI Get started synthesizing text-to-speech with our command-line tool.

Glossary

Term Definition
Utterance A unit of input for Octave. Contains text, voice, description, speed, and trailing_silence.
Generation The total generated audio output, referenced by generation_id.
Snippet A segment of the total generated audio output, referenced by snippet_id.

Streaming and non-streaming

The TTS API supports both streaming and non-streaming (synchronous) responses.

Streaming endpoints return audio as it is generated so playback can begin quickly, while non-streaming endpoints return the full result after processing completes.

Mode Direction Endpoints Typical use cases
Streaming (HTTP) Output only /v0/tts/stream/json,
/v0/tts/stream/file
Real-time playback, low perceived latency, pipelines that process chunks.
Streaming (WebSocket) Input & output /v0/tts/stream/input Interactive UIs that send text incrementally and receive continuous audio.
Non-streaming Single response /v0/tts,
/v0/tts/file
Simple integrations, saving files, predictable end-to-end timing.

Unidirectional streaming (HTTP)

Emits a sequence of JSON objects, each including a base64 audio and metadata.

Sends a continuous stream of raw audio bytes (for example audio/mpeg).

Bidirectional streaming (WebSocket)

Send text incrementally and receive audio continuously over the same connection.

Non-streaming (HTTP)

Returns a JSON payload with the entire audio as a base64 string.

Returns a downloadable audio file such as audio/mpeg.

Choosing which response type

Ultra low latency streaming: instant mode

Instant mode is a low-latency streaming mode designed for real-time applications where audio playback should begin as quickly as possible. Unlike standard streaming—which introduces a brief lead time before the first audio chunk is sent—instant mode begins streaming audio as soon as generation starts. Instant mode is enabled by default.

How instant mode works

Instant mode does not change the format of streamed responses—each chunk includes the same metadata; however chunks in instant mode will be smaller and begin to arrive more quickly.

Enabling/disabling instant mode

When to disable instant mode

Developer tools

Hume provides a suite of developer tools for integrating TTS.

API Reference See our API reference for TTS streaming and non-streaming endpoints. SDKs Open source SDKs for streaming and non-streaming. Stream audio, handle files, and integrate quickly. CLI A command-line tool that allows direct interaction with Hume’s TTS API, ideal for testing, automation, and rapid prototyping. MCP Server Run the Hume’s TTS MCP server to expose TTS tools to compatible clients. Sample code Open source examples you can copy, run, and adapt to get started quickly.

API limits

The following limits apply to Hume’s Text-to-Speech API.

Limit Value
Request rate limit (HTTP) Defined by your subscription tier
Maximum text length 5,000 characters per Utterance
Maximum description length 1,000 characters per Utterance
Maximum generations per request 5
Supported audio formats MP3, WAV, PCM