agentsvc.io

services / tts

textoperational · 1516 ms

Text to Speech

Convert text to natural speech audio (MP3, WAV or OGG) with neural Piper voices: en-female, en-male (US English) and de-male (German); all voices public domain or CC0, so the audio can be used commercially. Params: text (required, max 2000 characters), voice (default en-female, or picked from language), language (en or de), speed (0.5 to 2, default 1), format (mp3 default, wav, ogg). Returns audio_base64 and duration.

Run free trial ↗1 free call per day with the example input. Paid: $0.010 USDC, no limit.

Call it

import { wrapFetchWithPayment, x402Client } from "@x402/fetch";
import { registerExactEvmScheme } from "@x402/evm/exact/client";
import { privateKeyToAccount } from "viem/accounts";

const client = new x402Client();
registerExactEvmScheme(client, { signer: privateKeyToAccount(process.env.EVM_PRIVATE_KEY) });
const payFetch = wrapFetchWithPayment(fetch, client);

const res = await payFetch("https://agentsvc.io/api/v1/proxy/tts", {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({"text":"Hello from agentsvc. Your report is ready.","voice":"en-female"}),
});
const { data, payment } = await res.json();  // payment.transaction = on-chain receipt

Input

FieldTypeDescription
text *string
voiceen-female | en-male | de-male
languageen | de
speednumber = 1
formatmp3 | wav | ogg = "mp3"

Output

FieldTypeDescription
formatstring
mime_typestring
duration_secondsnumber
size_bytesinteger
voicestring
languagestring
genderstring
voice_licensestring
charactersinteger
audio_base64string

Example response (data)

{
  "format": "mp3",
  "mime_type": "audio/mpeg",
  "duration_seconds": 2.93,
  "size_bytes": 23658,
  "voice": "en-female",
  "language": "en-US",
  "gender": "female",
  "voice_license": "public domain",
  "characters": 42,
  "audio_base64": "SUQzBAAAAAAAI1RTU0UAAAAPAAADTGF2…"
}