Render text to speech with timing information and subtitles

This method takes text input and generates an audio clip, along with word timing data. Word timing output includes word level timing in a json file, an srt file, and vtt file.

Recent Requests
Log in to see full request history
TimeStatusUser Agent
Retrieving recent requests…
LoadingLoading…
Body Params
integer
enum
required
string
required
length between 1 and 1000
string
enum

Which model to use. Note that not all voices are available on all models.

Allowed:
library_ids
array of strings

List of ids of replacement libraries to use when processing

library_ids
audio_configs
object
Headers
boolean

Enables limited SSML translation for input text

string
enum
Defaults to application/json

Generated from available response content types

Allowed:
Responses

Language
Credentials
Header
LoadingLoading…
Response
Click Try It! to start a request and see the response here! Or choose an example:
application/zip
application/json