# Everlit TTS API

Version: 1.1.0
Last Updated: 2026-09-01

High-quality text-to-speech as a simple REST API: send text, get back a
production-ready MP3 narrated by any voice in your catalog — including your
own cloned voices — with word-level timestamps included.

**Base URL:** `https://api.everlit.audio`

---

## Quickstart

```bash
# 1. Create a speech job
curl -s -X POST "https://api.everlit.audio/v1/speech" \
  -H "Authorization: Bearer api_YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"text": "Hello from Everlit.", "voice": "109"}'
# => 202
# {
#   "id": "spch_01jx...", "status": "queued",
#   "audio_url": "https://api.everlit.audio/v1/speech/spch_01jx.../audio", ...
# }

# 2. Poll until status is "succeeded" (typically 15-60s), then download.
#    audio_url is known from the moment you submit and never changes.
curl -sL -o hello.mp3 \
  -H "Authorization: Bearer api_YOUR_KEY" \
  "https://api.everlit.audio/v1/speech/spch_01jx.../audio"
```

Prefer not to poll? Pass a `callback_url` and we POST you the finished job —
see [Webhooks](#webhooks).

---

## Client examples

End to end in three languages: create a job, poll every few seconds, download
the MP3. All three use only the standard HTTP client; there is no SDK to
install. Replace `109` with a voice id from `GET /v1/voices`.

```python
import os, time, requests

BASE = "https://api.everlit.audio"
H = {"Authorization": f"Bearer {os.environ['EVERLIT_API_KEY']}"}

job = requests.post(f"{BASE}/v1/speech", headers=H,
                    json={"text": "Hello from Everlit.", "voice": "109"}).json()

while job["status"] in ("queued", "processing"):
    time.sleep(3)
    job = requests.get(f"{BASE}/v1/speech/{job['id']}", headers=H).json()

if job["status"] != "succeeded":
    raise SystemExit(f"job failed: {job['error']}")

# audio_url answers 302 -> storage; requests follows it and drops the bearer.
with open("hello.mp3", "wb") as f:
    f.write(requests.get(job["audio_url"], headers=H).content)
```

```javascript
// Node 18+ (global fetch)
const BASE = "https://api.everlit.audio";
const H = { Authorization: `Bearer ${process.env.EVERLIT_API_KEY}` };

let res = await fetch(`${BASE}/v1/speech`, {
  method: "POST",
  headers: { ...H, "Content-Type": "application/json" },
  body: JSON.stringify({ text: "Hello from Everlit.", voice: "109" }),
});
let job = await res.json();

while (job.status === "queued" || job.status === "processing") {
  await new Promise((r) => setTimeout(r, 3000));
  job = await (await fetch(`${BASE}/v1/speech/${job.id}`, { headers: H })).json();
}
if (job.status !== "succeeded") throw new Error(`job failed: ${JSON.stringify(job.error)}`);

// fetch follows the 302 and drops the bearer on the cross-origin hop.
const audio = await fetch(job.audio_url, { headers: H });
await require("fs/promises").writeFile("hello.mp3", Buffer.from(await audio.arrayBuffer()));
```

```ruby
require "net/http"
require "json"

BASE = URI("https://api.everlit.audio")
AUTH = { "Authorization" => "Bearer #{ENV.fetch('EVERLIT_API_KEY')}" }

http = Net::HTTP.new(BASE.host, 443).tap { |h| h.use_ssl = true }
res = http.post("/v1/speech", { text: "Hello from Everlit.", voice: "109" }.to_json,
                AUTH.merge("Content-Type" => "application/json"))
job = JSON.parse(res.body)

while %w[queued processing].include?(job["status"])
  sleep 3
  job = JSON.parse(http.get("/v1/speech/#{job['id']}", AUTH).body)
end
abort "job failed: #{job['error']}" unless job["status"] == "succeeded"

# Net::HTTP does not follow redirects: take the 302 by hand, and send the
# storage request WITHOUT the bearer (the URL is self-authenticating).
redirect = http.get(URI(job["audio_url"]).path, AUTH)
File.binwrite("hello.mp3", Net::HTTP.get(URI(redirect["location"])))
```

The MCP server needs none of this — an agent calls `create_speech`, then
`get_speech` until `status` is `succeeded`. See [MCP](#mcp).

---

## Authentication

Every request needs an `api_` key as a bearer token:

```
Authorization: Bearer api_...
```

Keys are minted in Studio under **API → Keys**. They are server-side
credentials — never embed one in a browser or mobile app. The API is
pay-as-you-go: if you see `tts_api_not_enabled`, add a card under
**Billing → TTS API** first (see [Billing](#billing)).

**The full key is shown exactly once, when it is minted.** We store only a
SHA-256 digest and the first 12 characters, so it cannot be retrieved later —
copy it into your secret store at creation time. If a key is lost, revoke it
and mint a new one; the key list identifies each one by its prefix and label.

---

## Endpoints

| Method | Path | Purpose |
|---|---|---|
| `POST` | `/v1/speech` | Create a synthesis job → `202` |
| `GET` | `/v1/speech/{id}` | Fetch a job (poll this) |
| `GET` | `/v1/speech` | List jobs, newest first (`limit`, `starting_after`, `status`, `q`) |
| `GET` | `/v1/speech/{id}/audio` | Download the audio (302 redirect) |
| `GET` | `/v1/speech/{id}/timings` | Word-level timestamps (JSON) |
| `DELETE` | `/v1/speech/{id}` | Delete the audio → `204` |
| `GET` | `/v1/voices` | Voices available to your account |
| `GET` | `/v1/languages` | Supported languages, billing multipliers, transcript availability |
| `GET` | `/v1/usage` | Usage aggregates |
| `POST` | `/v1/mcp` | MCP server (Model Context Protocol) for AI agents — see [MCP](#mcp) |

## Create speech

`POST /v1/speech`

| Parameter | Type | Required | Description |
|---|---|---|---|
| `text` | string | yes | Text to narrate, up to your plan's per-request cap. Markup and control sequences are stripped. |
| `voice` | string | yes | A voice `id` from `GET /v1/voices`. |
| `language` | string | no | ISO 639-1 hint (e.g. `"en"`, `"es"`) — see `GET /v1/languages`. Defaults to auto-detection. Never lowers the billing rate (it only disambiguates Japanese vs Chinese readings of Han characters). |
| `format` | string | no | Only `"mp3"` today (44.1 kHz, 128 kbps). |
| `callback_url` | string | no | `https://` URL to POST the finished job to. |
| `metadata` | object | no | Up to 8 string values (512 bytes total), echoed back on the job. |

Unknown parameters are rejected with `400 unknown_parameter` — the API is
strict so that future additions are never breaking.

### The job resource

Returned by every endpoint and delivered to your callback — one shape:

```json
{
  "object": "speech",
  "id": "spch_01jx...",
  "status": "succeeded",
  "created_at": "2026-08-07T14:00:00Z",
  "completed_at": "2026-08-07T14:00:41Z",
  "expires_at": "2026-09-06T14:00:00Z",
  "voice": "109",
  "language": "en",
  "format": "mp3",
  "audio_url": "https://api.everlit.audio/v1/speech/spch_01jx.../audio",
  "timings_available": true,
  "usage": { "characters": 812, "billable_characters": 812, "audio_seconds": 58.3 },
  "audio": { "content_type": "audio/mpeg", "bytes": 934218, "duration_seconds": 58.3 },
  "error": null,
  "metadata": { "ref": "article-4512" }
}
```

`status` is one of `queued`, `processing`, `succeeded`, `failed`, `expired`,
`deleted`. Poll every 3–5 seconds; a typical job completes in 15–60 seconds,
long texts in a few minutes.

## Downloading audio

`audio_url` is **deterministic and stable**: it is returned the moment you
create the job, never changes, and works for the full 30-day retention
window. Authenticate with the same bearer key; the endpoint replies with a
302 redirect to short-lived storage, so `curl -L` and every HTTP client
follow it transparently. There are no expiring links to manage or cache.

While the job is still running the endpoint returns `409 job_not_complete`;
after retention or deletion it returns `410` (`audio_expired` /
`audio_deleted`).

**Retention:** audio and timings are kept for **30 days** from creation, then
deleted automatically. The job record itself survives (status `expired`) so
you can always reconcile usage against your bill. Download promptly or call
`DELETE /v1/speech/{id}` when you're done — deletion removes the audio but
never the usage record.

**Loudness:** output is voice-only and not run through a loudness-mastering
pass. Program loudness is typically around −16 LUFS but varies slightly by
voice; apply your own `loudnorm` downstream if you need broadcast compliance.

## Word timings

`GET /v1/speech/{id}/timings` — included with almost every job at no extra
cost (`timings_available` on the resource tells you):

```json
{
  "object": "speech.timings",
  "id": "spch_01jx...",
  "duration_seconds": 58.3,
  "words": [
    { "word": "Hello", "start": 0.0, "end": 0.31 },
    { "word": "from",  "start": 0.38, "end": 0.55 }
  ]
}
```

Times are seconds from the start of the audio file. Use them for karaoke-style
read-along highlighting, subtitles, or jumping playback to a sentence.

**The word list is your submitted text, verbatim** — words are taken from the
text you sent (whitespace-tokenized, punctuation preserved), never from a
speech recognizer's guess at it. Spoken expansions still map cleanly: "$5" is
one entry whose span covers the spoken "five dollars".

## Voices

`GET /v1/voices` returns every voice your key can use — Everlit's catalog
plus your own cloned voices (`"owned_by": "publisher"`). Filter with
`?language=en`. Voice `id`s are stable strings; pass them as `voice` when
creating speech.

---

## Languages

`GET /v1/languages` lists every supported language with two properties:

- **`billing_multiplier`** — see [Billing](#billing) below.
- **`transcripts`** — whether jobs in this language produce word-level
  timestamps (`/timings`). Currently 25 European languages; everywhere else
  jobs succeed with `timings_available: false`. Transcript coverage expands
  over time. The catalog itself covers every official EU language.

Launch catalog (69 languages):

- **×1** — English, Spanish, Portuguese, French, German, Italian, Dutch,
  Polish, Russian, Ukrainian, Czech, Slovak, Romanian, Greek, Hungarian,
  Finnish, Swedish, Danish, Bulgarian, Croatian, Estonian, Latvian,
  Lithuanian, Maltese, Slovenian, Irish, Norwegian, Catalan, Galician,
  Turkish, Vietnamese, Indonesian, Serbian, Bosnian, Belarusian, Kazakh,
  Mongolian, Uzbek, Georgian, Icelandic, Welsh, Basque, Esperanto, Swahili,
  Kinyarwanda, Ganda, Kabyle
- **×1.5** — Arabic, Hebrew, Persian, Urdu, Uighur, Kurdish (Sorani),
  Bengali, Hindi, Tamil, Telugu, Kannada, Marathi, Gujarati, Nepali,
  Malayalam, Punjabi, Odia
- **×2** — Korean, Thai
- **×2.5** — Japanese
- **×3** — Chinese, Cantonese

## Billing

You are billed on **`usage.billable_characters`: a script-weighted Unicode
character count of `text` exactly as you submitted it**, before any markup
stripping or text normalization.

**Why weights?** Our costs scale with audio length, and scripts differ in how
much speech one character buys: English yields roughly 1,050 characters per
spoken minute, Chinese roughly 300. So that every language costs about the
same **per finished hour of audio**, each character bills at its script's
rate: Latin/Cyrillic/Greek ×1, Arabic, Hebrew, and Indic scripts ×1.5,
Hangul and Thai ×2, kana ×2.5, Han ×3 (×2.5 when the text is Japanese — kana
present or `language: "ja"`). Mixed-script text is weighted per character;
`billable_characters = ceil(Σ per-character rates)`. The computation uses
standard Unicode script ranges, so your bill remains **reproducible from your
own request logs**; the `language` hint can never reduce it. Multipliers are
published on `GET /v1/languages` and only ever change in your favor for
already-supported languages.

Failed jobs are not billed. Repeating identical text is billed each
time — internal caching is our cost optimization, never a billing variable
(and never a way to get free synthesis).

### Pay-as-you-go

The API is prepaid and independent of your Everlit plan. In Studio under
**Billing → TTS API** you add a card and buy credit; each successful job is
priced from `usage.billable_characters` at the published per-million-character
rate (fractional cents — nothing is rounded up) and deducted from your balance
when it finishes. A request whose price
exceeds your balance is rejected up front with `402 insufficient_balance` —
nothing is queued and nothing is charged. Turn on **automatic top-up** (a
top-up amount and a balance threshold) and your card is charged whenever the
balance drops below the threshold, so production traffic never sees a `402`.
If you leave auto top-up off, we email your team when the balance falls below
your warning line (settable on the same page, $2 by default).

Operational limits (safety rails, not credits) for pay-as-you-go accounts:
**60 creates/min, 300 reads/min, 4 concurrent jobs, 20,000 characters per
request**. Need
more headroom? Ask support@everlit.audio — limits are set per account. The
per-request cap applies to raw text length. Your effective limits are returned
by `GET /v1/usage`; track consumption there or in Studio under **API → Usage**.

## Rate limits

Two independent per-key buckets: **creates** (`POST /v1/speech`,
`rate_limit_exceeded`) and **reads** (every other request, including each MCP
message, `read_rate_limit_exceeded`), so a polling loop can never block a
create. Every response carries `RateLimit-Limit`, `RateLimit-Remaining`, and
`RateLimit-Reset` (seconds) for the bucket it drew on. On `429` you also get
`Retry-After` — respect it, then retry with exponential backoff. Poll
`GET /v1/speech/{id}` every few seconds, not in a tight loop. A `503 capacity_unavailable` means
the problem is on our side; retry after the `Retry-After` value.

## Idempotency

Pass an `Idempotency-Key` header (any unique string, e.g. a UUID) on
`POST /v1/speech` to make retries safe: resubmitting the same key with the
same body returns the original job (`200` + `Idempotency-Replayed: true`
header) without charging you again. The same key with a *different* body is
rejected with `409 idempotency_key_reuse`. Keys are scoped to your account
and live for 30 days.

---

## Webhooks

Pass `callback_url` and we POST the full job resource to it when the job
reaches a terminal state, with header `X-Everlit-Event: speech.succeeded` or
`speech.failed`. Deliveries retry with backoff for ~30 minutes on non-2xx
responses. Only `https://` URLs are accepted.

Verify authenticity with `X-Everlit-Signature` and your webhook secret
(Studio → API → Usage):

```
X-Everlit-Signature: t=1754575200,v1=5257a869e7...
```

```python
import hmac, hashlib

def verify(payload_body, sig_header, secret):
    parts = dict(p.split("=", 1) for p in sig_header.split(","))
    expected = hmac.new(secret.encode(), f"{parts['t']}.{payload_body}".encode(),
                        hashlib.sha256).hexdigest()
    return hmac.compare_digest(expected, parts["v1"])
```

Reject deliveries whose timestamp `t` is more than 5 minutes old to prevent
replays.

---

## MCP

The API is also an [MCP](https://modelcontextprotocol.io) server, so Claude,
Cursor, Claude Code, and any other MCP-capable agent can narrate text
directly. It exposes the same operations as the REST endpoints as tools:

| Tool | Backs |
|---|---|
| `create_speech` | `POST /v1/speech` (accepts an `idempotency_key` argument) |
| `get_speech` | `GET /v1/speech/{id}` |
| `list_speech` | `GET /v1/speech` |
| `get_speech_timings` | `GET /v1/speech/{id}/timings` |
| `delete_speech` | `DELETE /v1/speech/{id}` |
| `list_voices` | `GET /v1/voices` |
| `list_languages` | `GET /v1/languages` |
| `get_usage` | `GET /v1/usage` |

**Endpoint:** `https://api.everlit.audio/v1/mcp` — Streamable HTTP transport,
stateless (no session ids), JSON responses. Authenticate every request with
the same `Authorization: Bearer api_...` header; there is no OAuth flow.

```json
{
  "mcpServers": {
    "everlit-tts": {
      "type": "http",
      "url": "https://api.everlit.audio/v1/mcp",
      "headers": { "Authorization": "Bearer api_YOUR_KEY" }
    }
  }
}
```

```bash
claude mcp add --transport http everlit-tts https://api.everlit.audio/v1/mcp \
  --header "Authorization: Bearer api_YOUR_KEY"
```

There is no blocking "wait" tool: while a job is `queued` or `processing`,
`create_speech` and `get_speech` results carry `poll_after_seconds` (20), and
the server instructions tell agents to wait that long between `get_speech`
calls. `get_speech` is not billed, but every MCP message counts against the
read rate limit (a `429` at the HTTP level, see [Rate limits](#rate-limits)).

Tool calls are metered, limited, and billed exactly like the REST endpoints —
they run the same code. A failed call returns an MCP tool error (`isError:
true`) whose `structuredContent.error` is the same envelope documented under
[Errors](#errors), plus `http_status` and, on `429`/`503`, `retry_after`
seconds. Audio itself is never returned inline: read `audio_url` from the
finished job and download it with your API key, as described under
[Downloading audio](#downloading-audio).

### Tools

Every tool takes a JSON object of arguments and returns the same JSON the
matching REST endpoint would, as `structuredContent` (and pretty-printed as
text). `create_speech` and `get_speech` add `poll_after_seconds` while the job
is still running.

<!-- mcp-tools:start (generated by `rake docs:tts_api`; do not edit by hand) -->

### `create_speech`

Start narrating text with a voice. Returns the speech job (status "queued") including its id and stable audio_url; poll get_speech (honoring poll_after_seconds) until status is "succeeded". Billed per submitted character.

| Argument | Type | Required | Description |
|---|---|---|---|
| `text` | string | yes | Text to narrate. Markup and control sequences are stripped. |
| `voice` | string or integer | yes | A voice id from list_voices. |
| `language` | string | no | Optional ISO 639-1 hint (e.g. "en", "es"); see list_languages. Defaults to auto-detection. |
| `format` | enum: mp3 | no | Output format. Only "mp3" today. |
| `callback_url` | string | no | Optional https:// URL to POST the finished job to. |
| `metadata` | object | no | Up to 8 string values (512 bytes total), echoed back on the job. |
| `idempotency_key` | string | no | Any unique string; resubmitting the same key and text returns the original job without a second charge. |

### `get_speech`

Fetch a speech job by id. status is one of queued, processing, succeeded, failed, expired, deleted. While queued/processing the result includes poll_after_seconds: wait that long before checking again.

| Argument | Type | Required | Description |
|---|---|---|---|
| `id` | string | yes | The job id (spch_...). |

### `list_speech`

List this account's speech jobs, newest first. Paginate with starting_after = the last id of the previous page while has_more is true.

| Argument | Type | Required | Description |
|---|---|---|---|
| `limit` | integer | no | Page size (default 20). |
| `starting_after` | string | no | Return jobs created before this job id. |
| `status` | enum: queued, processing, succeeded, failed, expired, deleted | no | Filter by status. |
| `q` | string | no | Free-text match on id, language, voice name, error code, or metadata. |

### `get_speech_timings`

Word-level timestamps for a succeeded job (when timings_available is true). Useful for captions and read-along highlighting.

| Argument | Type | Required | Description |
|---|---|---|---|
| `id` | string | yes | The job id (spch_...). |

### `delete_speech`

Permanently delete a job's audio and timings. The job record and its usage stay for billing reconciliation.

| Argument | Type | Required | Description |
|---|---|---|---|
| `id` | string | yes | The job id (spch_...). |

### `list_voices`

Voices available to this account (Everlit's catalog plus the account's own cloned voices). Use the id as create_speech's voice.

| Argument | Type | Required | Description |
|---|---|---|---|
| `language` | string | no | Filter by language prefix, e.g. "en" or "en-US". |
| `limit` | integer | no | Page size (default 100). |
| `starting_after` | string | no | Return voices with an id greater than this one. |

### `list_languages`

Supported languages with their billing multipliers and whether word timings are available.

_No arguments._

### `get_usage`

Character usage for this account: month to date, plan limits, and a daily breakdown for the window (default: this month).

| Argument | Type | Required | Description |
|---|---|---|---|
| `from` | string | no | ISO 8601 start of the window. |
| `to` | string | no | ISO 8601 end of the window. |

<!-- mcp-tools:end -->

---

## Errors

Errors use one envelope, with a stable machine-readable `code` — branch on
`code`, never on `message` text:

```json
{
  "error": {
    "type": "invalid_request_error",
    "code": "text_too_long",
    "message": "text is 34,102 characters; the maximum for your plan is 20,000.",
    "param": "text",
    "request_id": "ab12cd34"
  }
}
```

Include `request_id` in support tickets.

| HTTP | code | Meaning |
|---|---|---|
| 400 | `invalid_json` | Body is not a JSON object |
| 400 | `missing_parameter` | A required parameter is absent or the wrong type |
| 400 | `unknown_parameter` | Unrecognized parameter (check spelling) |
| 400 | `invalid_encoding` | `text` is not valid UTF-8 |
| 400 | `text_too_long` | Over your plan's per-request cap |
| 400 | `unsupported_language` | Not an accepted `language` code |
| 400 | `unsupported_format` | Only `mp3` today |
| 400 | `invalid_callback_url` | Must be https with a hostname |
| 401 | `missing_api_key` / `invalid_api_key` / `api_key_revoked` / `api_key_expired` | Credential problems |
| 402 | `insufficient_balance` | Prepaid balance doesn't cover the request — add credit |
| 403 | `wrong_key_type` | Use an `api_` key (not `pk_`/`tkn_`) |
| 403 | `voice_not_accessible` | Voice belongs to another account |
| 403 | `tts_api_not_enabled` | Pay-as-you-go not set up — add a card in Studio |
| 404 | `job_not_found` / `voice_not_found` / `timings_not_available` / `unknown_endpoint` | Nothing at that address |
| 409 | `idempotency_key_reuse` | Same key, different body |
| 409 | `job_not_complete` | Audio requested before the job finished |
| 410 | `audio_expired` / `audio_deleted` | Audio is gone (job record remains) |
| 422 | `no_speakable_text` | No letters or digits to narrate |
| 429 | `rate_limit_exceeded` / `read_rate_limit_exceeded` / `concurrency_limit_exceeded` | Slow down (see `Retry-After`) |
| 500 | `internal_error` | Our fault; the job was not billed |
| 503 | `capacity_unavailable` | Temporarily saturated; retry |

### Notes on input handling

- HTML/markup in `text` is stripped before narration but **counts toward
  billing** — send plain text.
- Bracketed Everlit control sequences (`[Everlit-...]`) are stripped.
- Control characters are removed; text is Unicode-normalized (NFC).

---

## Changelog

**1.1.0 — 2026-09-01**
- MCP server at `POST /v1/mcp` (Streamable HTTP, same API keys) with one tool per REST endpoint; running jobs report `poll_after_seconds`.
- Reads now have their own rate-limit bucket, separate from creates: new `429 read_rate_limit_exceeded`. `RateLimit-*` headers appear on every response.
- `401` responses carry a `WWW-Authenticate: Bearer` challenge.

**1.0.1 — 2026-08-31**
- `GET /v1/languages`: the supported-language catalog with billing multipliers and transcript availability.
- Script-weighted billing: `usage.billable_characters` (what you are billed) alongside `usage.characters` (what you sent). Latin-script text is unchanged.
- Optional `language` hint on `POST /v1/speech`.

**1.0.0 — 2026-08-07**
- Initial release.

---

## Support

- Email: support@everlit.audio
- OpenAPI spec: https://everlit.audio/docs/tts-openapi.yml
- Status: https://status.everlit.audio
