services / tts
textoperational · 1516 ms
Text to Speech
Convert text to natural speech audio (MP3, WAV or OGG) with neural Piper voices: en-female, en-male (US English) and de-male (German); all voices public domain or CC0, so the audio can be used commercially. Params: text (required, max 2000 characters), voice (default en-female, or picked from language), language (en or de), speed (0.5 to 2, default 1), format (mp3 default, wav, ogg). Returns audio_base64 and duration.
Run free trial ↗1 free call per day with the example input. Paid: $0.010 USDC, no limit.
Call it
Input
| Field | Type | Description |
|---|---|---|
| text * | string | |
| voice | en-female | en-male | de-male | |
| language | en | de | |
| speed | number = 1 | |
| format | mp3 | wav | ogg = "mp3" |
Output
| Field | Type | Description |
|---|---|---|
| format | string | |
| mime_type | string | |
| duration_seconds | number | |
| size_bytes | integer | |
| voice | string | |
| language | string | |
| gender | string | |
| voice_license | string | |
| characters | integer | |
| audio_base64 | string |
Example response (data)
{
"format": "mp3",
"mime_type": "audio/mpeg",
"duration_seconds": 2.93,
"size_bytes": 23658,
"voice": "en-female",
"language": "en-US",
"gender": "female",
"voice_license": "public domain",
"characters": 42,
"audio_base64": "SUQzBAAAAAAAI1RTU0UAAAAPAAADTGF2…"
}