Skip to main content

Audio APIs

Sunset Notice: Melo TTS is being discontinued. This model will be removed in a future update. Please plan your migration accordingly.
Convert text to natural-sounding speech using Melo TTS. Generate high-quality audio in multiple languages with customizable speed and speaker options.

Text-to-Speech

Endpoint

Basic Example

Request Parameters

Required Parameters

Optional Parameters

Response Format

The API returns a JSON object containing base64-encoded MP3 audio:

Decoding the Response

Supported Languages

Melo TTS supports 6 languages with various speaker options:
English has multiple speaker variants for different accents: US (American), BR (British), INDIA (Indian), and AU (Australian).

Using Different Languages and Speakers

Multilingual Example

Speed Control

Adjust the speed parameter to control how fast the speech is generated:

Example

Use slower speeds (0.5-0.8) for instructional content or accessibility needs. Use faster speeds (1.2-1.5) for content review or when listeners prefer quicker playback.

Pricing

Rate: $5.00 per 1 million characters
There are no character limits per request. You are billed based on the total characters processed.

Use Cases

  • Voice assistants: Add natural speech to chatbots and virtual assistants
  • Audiobook generation: Convert written content to audio format
  • Accessibility: Make content accessible for visually impaired users
  • Video narration: Generate voiceovers for videos and presentations
  • Language learning: Create pronunciation examples in multiple languages
  • Notification systems: Generate audio alerts and announcements

Next Steps

Text APIs

Generate text with large language models

Vision Language Models

Analyze images with multimodal AI

Image APIs

Generate images from text prompts