How-to · 3 steps

How to give an agent a voice
with Monid.

Give this to your agent
$set up https://monid.ai/SKILL.md, then say this sentence out loud for me and show the cost of each step

and let it take it from there.

›say this out loud, then read it back
#JobWhat came backCost
1Speak ita hosted link, not bytes
2The faster modelsame words, half the bill
3Read it backthe brand name came back wrong
4Tell it the wordsometimes fixes, sometimes not
5Two voicesone file, 6.4 seconds
6Ten seconds of musicexactly ten, billed exactly
7A sound bedfive seconds, deterministic
8One file, four endpointsfour different billed durations
9The other vendorflat per song, no duration
10Total19 runs
one sentence · out and back for $0.003756 · and it came back wrongagent running$0.069475

Real run, 2026-09-25. Seven MP3s downloaded and measured.

Step 1

Set up Monid.

One line. It installs the CLI and asks for an API key from app.monid.ai. New accounts start with $1.00.

Say this to your agent
>set up https://monid.ai/SKILL.md
Step 2

Pick what you need. Say it in one sentence.

Click a job. Paste the sentence, fill in the brackets.

Turn text into speech

Say this to your agent
>say <text> out loud with elevenlabs on monid. it returns a hosted link rather than audio bytes, so download the file in the same run: the link is documented to last an hour and the file a week
What the agent runsdone
elevenlabs /v1/text-to-speech$0.05 / 1k chars
elevenlabs /v1/text-to-dialogue$0.10 / 1k chars
elevenlabs /v1/voicesread-only
minimax /v1/t2a_v2catalog rate
What came backone sentence, round trip · 2026-09-25
The same sentence, two modelsUS dollars for 69 characters
the default model$0.0069
the faster model$0.00345
what comes backa download link on Monid's own file service, never inline bytes, with a documented one-hour link expiry and a seven-day file expiry
asking again is safefetching the run a second time mints a fresh link and the old one keeps working, so a link that has expired costs nothing to replace
but the file clock does not restartthe seven days run from the original call, not from the refetch, so a fresh link on day seven buys you hours rather than another week
and there IS a pollan earlier version of this page said there is no polling and no job id anywhere here, which contradicts the platform's own documented run-then-poll loop; fetching the run by id IS that poll
one field halves it, and you must pass itnaming the faster model cut the bill exactly in half on the same words on both runs, and it produced slightly more audio, so the per-minute gap is a little wider than 2x. Left unset, the default is the dearer model, so the saving is something you ask for rather than something you get
the default voice is not in the listthe documented default voice id does not appear among the voices the voice endpoint returns, so pick one from the list rather than relying on the default. The list is account-scoped and came back with 42 and then 44 entries, so read it rather than quoting a count
$0.006969 characters spoken

Read it back and check it

Say this to your agent
>transcribe the audio you just made with elevenlabs on monid, then diff the transcript against the text you sent character by character. expect a word with a diacritic or an unusual spelling to come back wrong and expect the final full stop to be missing. try that word in keyterms, but check the result: on our second run keyterms turned a wrong word into a non-word rather than fixing it
What the agent runsdone
elevenlabs /v1/speech-to-text$0.22 / hour
elevenlabs /v1/forced-alignment$0.22 / hour
elevenlabs /v1/text-to-speech$0.05 / 1k chars
What came backone sentence, round trip · 2026-09-25
it came back wrongon a clean machine-generated clip with no noise, no accent and no crosstalk, a word came back as a different word and the closing full stop was dropped; that dropped full stop reproduced on a second sentence in all three runs
what breaks is not always the brand nameon a second sentence the brand name and the unusual surname both survived and the word with a diacritic did not: Wroclaw came back as Rockwool
and keyterms does not reliably fix itthe first run it did. On the second it made it worse, turning the wrong word into a non-word, at the same $0.000069 extra. So treat keyterms as worth a try and then check, not as the fix
what that impliesif a synthetic clip mis-transcribes a proper noun, a real recording will too; tell the transcriber the names it should expect rather than checking afterwards
the round trip is cheaptext out loud and back into text cost $0.003756 in total, which is the cheapest verification step on this page
$0.003756the whole loop

Making audio against reading it

Say this to your agent
>before you plan a job, measure what a minute of audio costs to make and what a minute costs to read, on the models you will actually use: characters per minute is not a constant, and the model this command picks by default is the expensive one. then multiply both out to a thousand minutes
What the agent runsdone
elevenlabs /v1/text-to-speech$0.05 / 1k chars
elevenlabs /v1/speech-to-text$0.22 / hour
elevenlabs /v1/music$0.15 / minute
What came backone sentence, round trip · 2026-09-25
One minute of audioUS dollars a minute
compose music$0.15
speak it$0.0469
transcribe it$0.003667
the multiplereading is roughly twelve to forty times cheaper than making, depending which model and which route; on defaults it came out at about 27x
the rate is not a constant, so measure itcharacters per minute of delivered audio came out at 938 on the first sentence and at 869 and 976 on a second, for the two models, so the same text buys a different number of minutes depending on the voice
and the default model is the dear onethe halving only happens if you pass the model explicitly; left to the default, a thousand minutes of speech is about $97 rather than about $43
music has two shapes, not oneone provider bills per minute and the other bills a flat price per song regardless of length, which came out at about $57 a thousand minutes against about $150, so the cheaper composer depends entirely on how long your pieces are
per character ranks differentlythe two-voice route costs the same per character as the single-voice one and less per minute, because two speakers with a laugh in them talk more slowly
$0.00arithmetic on measured rates

Music and a sound bed

Say this to your agent
>make me <n> seconds of <mood> music with elevenlabs on monid, and a <n> second sound bed. this provider bills exactly the duration you ask for, so ask for exactly what you need. if you use the other music provider instead, note that it takes no duration parameter at all and bills a flat price per song, so short pieces are dear there and long ones are cheap
What the agent runsdone
elevenlabs /v1/music$0.15 / minute
elevenlabs /v1/sound-generation$0.12 / minute
minimax /v1/music_generationcatalog rate
What came backone sentence, round trip · 2026-09-25
it is deterministicyou set the duration and it bills exactly that duration: $0.025 for ten seconds of music and $0.01 for a five-second bed, with no rounding surprise
so budget in secondsunlike the character-priced routes, there is nothing to measure afterwards; the number you ask for is the number you pay
the second vendor works and prices differentlyan earlier version of this page wrote it off as dead. On a second run it answered and billed a flat fifteen cents for a song of about two and a half minutes, with no duration parameter to set
a third one is the dead onea different provider's music endpoint, flagged unknown in the catalog rather than stable, timed out with a 504; it billed nothing, as every genuine failure here did
so ask for exactly what you need is provider-specificit is true of the per-minute provider and meaningless on the flat-rate one
$0.035ten seconds and five

One file, three different durations

Say this to your agent
>if you send the same audio to more than one endpoint, record what duration each one billed: on our runs one identical file was billed at four different durations by four endpoints on the same provider, spreading about fifty percent. one of them also enforces a 4.6 second minimum that is in no schema, and the CLI hides that rejection behind a generic 400 with no output file, so read the run record rather than the client error
What the agent runsdone
elevenlabs /v1/speech-to-text$0.22 / hour
elevenlabs /v1/forced-alignment$0.22 / hour
elevenlabs /v1/speech-to-speech$0.12 / minute
elevenlabs /v1/audio-isolation$0.12 / minute
What came backone sentence, round trip · 2026-09-25
The same 4.133 second file, billedseconds charged
transcription6 seconds
forced alignment5 seconds
isolation4 seconds
voice conversion4 seconds
the spreadone file, four endpoints on the same provider, four different billed durations: 67% apart on the first file and about 50% on a second, so the shape holds and the percentage does not
and the floor contradicts itselfthe endpoint that refuses a 4.55 second file for being under its 4.6 second minimum billed a 5.108 second file as four seconds, which is below the minimum it just enforced
why it matters at scalea pipeline that estimates cost from the file's real duration will be wrong in three different directions depending on which endpoint it used
the undocumented floorone endpoint enforces a 4.6-second minimum that appears nowhere in its schema, and rejects anything shorter
and the CLI hid itthat rejection printed only a generic HTTP 400 and wrote no output file, while still creating three real runs whose records carried the actual too-short message
$0.00the spread, not the cost

What this aisle cannot do

Say this to your agent
>tell me plainly what is not here before i design around it: there is no dubbing, and nothing in these three aisles streams a generated file back to me. do NOT conclude from that that live voice is impossible: a live phone call is metered by the second elsewhere on the platform, and the run stays open for the length of the call. check whether a cloned voice can be USED before you tell me one cannot be MADE: they are different questions
What the agent runsdone
elevenlabs /v1/voicesread-only
elevenlabs /v1/speech-to-speech$0.12 / minute
minimax /v1/t2a_v2catalog rate
What came backone sentence, round trip · 2026-09-25
no voice CREATION, which is not the same as no cloned voicesnothing in these aisles creates or clones a voice. But one provider's own voice setting states that system, cloned and AI-designed voices are all accepted, and it carries a timbre control that blends four voices into a new one, so a cloned voice made elsewhere can be used here
and the voice list is not a fixed 42it returned 44 on a second run and it is scoped to the account, so it is a list to read rather than a number to quote
no streaming, and that does NOT rule out live voiceno endpoint in these three aisles streams a generated file back to you; every route is request, wait, download. An earlier version of this page concluded from that that a real-time conversational agent is impossible. It is not: a phone call elsewhere on the platform is metered per second at about thirty cents a minute, the run stays open for the length of the call and stopping the run hangs up
a stream flag exists but is pinned offtwo endpoints do declare a stream parameter, fixed at false, so an agent grepping the schemas will find one and must read the constant rather than the key
no dubbingthere is a voice-to-voice conversion endpoint, and it is not the same thing as translating and re-voicing a track
read the notes, not the summaryfive of nine summaries describe the output as base64 audio and all five returned a hosted link instead; the notes block was the accurate one, and it documents base64 as a fallback for when the file cannot be saved rather than as something that never happens
$0.00know the holes first
Step 3

Take the cheapest route with Monid.

Reading audio is more than ten times cheaper than making it. Speak once, verify with the cheap half, and tell the transcriber the names it should expect.

Pick the voice · $0.0000
List the voices$0.00read-only, account-scoped
Check the default$0.00it is not in the list
Make it · cheapest first
Fast model$0.0034
Default model$0.0069
Two voices$0.0072
Sound bed$0.01
Music$0.025
Read it back
Transcribe it$0.0003
Diff the text$0.00check the proper nouns
Back to you
rowspoken bycost
1Fast model$0.0038
2Fast model$0.0038
3Two voices$0.0179
4Fast model$0.0038
5Music$0.0529
5 of 5 rows$0.0821
Say this to your agent
>say each of these out loud, then read them back and tell me which ones came back wrong: <paste lines>
“read it back to check”a clean synthetic clip mis-transcribed a proper noun, and verifying cost a thousandth of a cent
“try keyterms, then check”the same term fixed a wrong word on one run and turned it into a non-word on another, at the same $0.000069
“download in the same run”speech comes back as a link with a documented one-hour expiry, not as bytes
“pick a voice from the list”the documented default voice id is not in the list the voice endpoint returns, and that list is account-scoped

One key. 1,700+ tools.

Eleven endpoints across three aisles, and today nine of them are one vendor.

Text into speechelevenlabs · $0.05 / 1k chars
Two voices, one fileelevenlabs · $0.10 / 1k chars
Speech back into textelevenlabs · $0.22 / hour
Word-level alignmentelevenlabs · $0.22 / hour
Compose musicelevenlabs · $0.15 / minute
Sound effects and bedselevenlabs · $0.12 / minute
Isolate a voice from noiseelevenlabs · $0.12 / minute
A second vendor, listedminimax · catalog rate

+ 1,700 more across 55 providers

Browse the catalog →