No endpoints here yet
Check back soon — the catalog grows weekly.
Try it with your agents
01
Generate campaign assets end to end
Set up https://monid.ai/SKILL.md, and then use Monid to generate a product hero image, a 15-second teaser video, and a voiceover for a smart espresso machine launch.02
Generate game-ready 3D assets
Set up https://monid.ai/SKILL.md, and then use Monid to generate a 3D model and matching materials for a low-poly medieval blacksmith forge and hand me the asset files.How-to guides
What does a video clip cost to generate?read the pricing unit before the price
Step 1
Set up Monid.
One line. It installs the CLI and asks for an API key from app.monid.ai. New accounts start with $1.00.
Say this to your agent
>set up https://monid.ai/SKILL.md
Step 2
Pick what you need. Say it in one sentence.
Click a job. Paste the sentence, fill in the brackets.
The cheapest clip that works
Say this to your agent
>generate a five second clip of <shot> with minimax h3-max-turbo at 480P on monid, passing the parameters in the request BODY: every video endpoint here takes a body, and sent as query params none of them run at all. tell me the billed seconds so i can check the price against the rate
What the agent runsdone
minimax /v1/video/minimax-h3-max-turbo$0.025 / s
minimax /v1/video/minimax-h3 · not run$0.038 / s
kling /v1/video/kling-2.5-turbo-t2v$0.042 / s
gemini /v1/video/omni-flash-t2v$0.04 / s
alibaba /v1/video/wan3.0 · not run$0.05 / sWhat came backfive real clips · 2026-09-24
The same five-second clipUS dollars
cheapest$0.125
with audio$0.20
mid tier$0.21
4K tier$2.10
what came backa playable 480P clip, five billed seconds, in fifteen seconds of wall clock
the spreadthe identical five-second prompt is $0.125 here and over $1.40 on the top audio tier
the health flagthe cheapest endpoint is marked degraded in the catalog and still answered first try
plain inspect crashes on four of these modelsit dies mid-pricing with a type error and prints no input schema at all, so an agent told to inspect before running stalls on a model that works; ask for JSON output instead and it answers
a dead job can still say COMPLETEDtwo runs came back with a run status of COMPLETED over a provider response of 500 and a body reading failed. They correctly billed nothing, but anything branching on the run status reports a success that has no file behind it
the rulepick the tier by what the shot needs, because the model choice is an eleven-fold decision
$0.125five seconds at $0.025
What a second actually costs
Say this to your agent
>run the same prompt again at ten seconds on the same model and show me both bills side by side, so i can see whether the price is per second or per call
What the agent runsdone
minimax /v1/video/minimax-h3-max-turbo$0.025 / s
minimax /v1/video/minimax-hailuo-2.3 · not run$0.28 / bundle
gemini /v1/video/omni-flash-t2v$0.04 / sWhat came backfive real clips · 2026-09-24
Double the durationUS dollars
5 seconds$0.125
10 seconds$0.25
the proofsame model, same prompt, twice the duration, exactly twice the bill
why it mattersa per-second rate means a thirty-second cut is six times a five-second one, not a little more
the exceptionone model is sold as fixed bundles, so six seconds is the floor and five is not offered
the auto trapsetting duration to auto pre-holds a full thirty seconds of charge, and it does it on three of the models here rather than one, so never leave the duration unset
$0.25ten seconds at $0.025
Turn a still into a shot
Say this to your agent
>animate <image url> into a five second clip with kling 2.6 image to video at 720p on monid, audio off, passing the parameters in the request BODY rather than as query params. use an image host that serves requests with no user agent: wikimedia returns 403 to the provider's fetcher and the failure takes a minute to surface as an unhelpful message about getting the contents of the file
What the agent runsdone
kling /v1/video/kling-2.6-i2v$0.042 / s
minimax /v1/video/minimax-h3-max-turbo$0.025 / s
gemini /v1/video/omni-flash-i2v · not run$0.04 / s
alibaba /v1/video/wan2.7-i2v · not run$0.10 / s
kling /v1/video/kling-3.0-i2v · not run$0.084 / sWhat came backfive real clips · 2026-09-24
what came backa clip from a single still, in 54 to 64 seconds of wall clock across two runs
the image host is the hard parttwo runs failed before one worked: the most obvious public image host returns 403 to any client sending no user agent, which is what the provider's server-side fetcher is, and the error it produces names the file rather than the refusal
and 720p is not what arrivesa request for 720p came back 1172 by 784, so the resolution parameter sets the price tier rather than the geometry: measure the file, do not trust the request
the roundingthe output ran 5.041 seconds and the bill was for five, because it bills whole requested seconds
the cheap alternativethe cheapest model also does image to video with first and last frame, at 480P for $0.025 a second
the 4K jumpthe same job on the top tier at 4K is $0.42 a second, so five seconds is $2.10
$0.21five seconds at $0.042
A clip that comes with sound
Say this to your agent
>generate <shot> with gemini omni-flash on monid and include the audio cue in the prompt. tell me whether the rate changes with duration
What the agent runsdone
gemini /v1/video/omni-flash-t2v$0.04 / s
kling /v1/video/kling-2.6-t2v · not run$0.14 / s with audio
kling /v1/video/kling-3.0-turbo-t2v · not run$0.112 / s
elevenlabs /v1/text-to-speech · not runto $0.22 / callWhat came backfive real clips · 2026-09-24
what came backa 360p clip with speech, music and ambience generated together, in 34 seconds
flat is the norm, not the exceptionan earlier version of this page called this the only model with one rate across every duration. It is the other way round: nearly every endpoint here is flat per second, and the bundle-priced one is the lone exception
and it is not the cheapest way to get soundanother model does audio at 720P for $0.10 a second, and a turbo tier does it at $0.112, so the $0.14 1080p tier this page used to name as the cheapest alternative is neither the cheapest nor the only one
the link trapthe signed download link expires in one hour while the file lives seven days: fetch it, do not save the url
$0.20five seconds at $0.04
When the premium tier is worth it
Say this to your agent
>price a five second clip on the premium token-billed tier before running anything, and tell me what it would cost against the cheapest model. confirm with me before you run it
What the agent runsdone
bytedance seedance-2.5 · not run$10.70 / 1M tokens
kling /v1/video/kling-3.0-i2v · not run$0.42 / s at 4K
alibaba /v1/video/wan2.7-videoedit · not run$0.10 / s, both waysWhat it returnsshape of the answer
the pricea five second clip is $1.16 at 720p, which is its default and also its ceiling, or $0.50 at 480p: about one and a half times the four real clips on this page put together
the number this page used to printan earlier version said $3.50 a clip. That figure is not a clip price at all: it is the per-million-token RATE of the CHEAPEST tier, which the catalog prints as $3.5 / 1M. Read the unit next to a token price before you turn it into a clip price
the unitit bills by tokens, but duration moves the bill far more than resolution does: across this tier resolution spans 2.3x and duration spans 7.5x
the hidden route, with a caveatone aggregator lists a flat $0.44 per call while its model field can route the job elsewhere; on our second look its documented models were the previous generation rather than this tier, so read the model list before you trust the flat price
the editing trapone edit endpoint bills the input seconds as well as the output, so a five second edit costs like ten
$1.16five seconds at 720p · catalog rate
Read more
View all →
MiniMax H3 Max on Monid
The model behind the cheapest seconds here, and what it can and cannot frame.

The best social media scraping API in 2026
Before you generate a shot, seeing what already works is the cheaper call.

LLM gateway or MCP gateway
Why one key across every model matters more once the models bill by the second.
What does an image or a 3D mesh cost to generate?price one run, because the ranking reverses between tiers
Step 1
Set up Monid.
One line. It installs the CLI and asks for an API key from app.monid.ai. New accounts start with $1.00.
Say this to your agent
>set up https://monid.ai/SKILL.md
Step 2
Pick what you need. Say it in one sentence.
Click a job. Paste the sentence, fill in the brackets.
What an image actually costs
Say this to your agent
>price one image before you plan a batch: on a token-billed route the bill is dominated by the output-token rate, so read the run's own cost rather than reading the input tier and stopping there. and do not sort the catalog by price without checking: one vendor's tiered entries list an amount of zero while billing three cents an image
What the agent runsdone
alibaba /qwen-image$0.030
minimax /image$0.0035What came backone prompt, four sizes · 2026-09-25
What the listing says against what it billedUS dollars for one image
what the run billed$0.006025
what the listing implies$0.000145
the gap is the tier you read, not a missing ratethe output-token rate of $0.03 per thousand is 97% of what an image costs, and it IS in the catalog next to the input tiers. An earlier version of this page said the listing omitted it. What is inspect-only is the worked low, medium and high examples
so budget from a runthe input tier alone is out by a factor of forty on this route, so one call and its own cost field is the reliable estimate
and a cheapest-first sort can point at the dearestone vendor's tiered entries list a price amount of zero in the catalog while billing three cents an image, which puts a vendor 8.6x dearer at the top of a price sort
the cheapest vendor is not downan earlier version of this page said the lowest-listed vendor 502s on every attempt. On a second run it answered first try and billed $0.0035, which is 1.7x below the floor this page quoted
status is not successthat 502 run reported its status as COMPLETED with a null output, so check the payload rather than the state
$0.006025one image, 1024 square
The same prompt on two vendors
Say this to your agent
>generate <prompt> at 1024 by 1024 on two vendors with monid, download both, and tell me the measured pixel dimensions and file sizes from the files themselves rather than from the response
What the agent runsdone
alibaba /qwen-image$0.030
alibaba /qwen-image-plus · not run$0.030What came backone prompt, four sizes · 2026-09-25
Same prompt, same 1024 squareUS dollars, and bytes delivered
vendor A · 1,906,937 bytes$0.006025
vendor B · 1,438,205 bytes$0.030000
5.0x on pricethe same prompt at the same requested size cost $0.006025 on one vendor and $0.030000 on the other
and the cheaper one was bigger1,906,937 bytes against 1,438,205 between those two, so the price difference bought no more file. A third and cheaper vendor returned a 288 KB JPEG, so file size does not track price in either direction
the ranking flips at the MIDDLE tierthe flat-price vendor becomes 1.76x cheaper at the medium setting, not the high one. At high the token-billed vendor is about 7x DEARER, so the crossover is earlier and steeper than an earlier version of this page said
the dimensions were honestevery image across every size matched what was asked for on both runs, measured from the file header; note that a size of 4K resolves to a square 4096 rather than UHD, and that only one endpoint accepts that literal at all
$0.036two images, one prompt
Edit an image you already have
Say this to your agent
>edit <image> with openai on monid to <change>. expect the scene to be REGENERATED rather than the pixels preserved: the lighting and framing will shift even where the edit lands correctly, so compare the whole frame, not just the thing you named
What the agent runsdone
What came backone prompt, four sizes · 2026-09-25
what it didon a second image, asked to recolour a named object, it recoloured that object and added nothing. The regeneration is the reliable part: the lighting and framing moved even though the edit itself was correct
so the risk is drift, not substitutionan earlier version of this page said it adds a new object beside the one you named. That did not reproduce; what reproduced on both runs is that nothing outside the edit is preserved
what it costs$0.014262, which is 2.37x the price of generating a fresh image on the same vendor
so when is it worth itwhen the composition is what you want to keep; if the whole scene is going to be regenerated anyway, a new prompt is cheaper and more direct
the prompt rewrite next doorthe other vendor extends your prompt by default. It is a documented parameter and the response tells you whether a rewrite happened; what it withholds is the rewritten TEXT, so you know it changed and not how
$0.0142622.37x a fresh image
A real 3D mesh
Say this to your agent
>generate a 3d mesh of <subject> with suzanne on monid, asking for glb, obj and stl together in one job rather than one format at a time. do NOT pass the face count the schema prints as its default: that value is rejected upstream and you get a 400. download inside fifteen minutes, or re-mint the link from the download endpoint, which is free
What the agent runsdone
suzanne /text-to-3d$0.80
suzanne /photo-to-3dcatalog rateWhat came backone prompt, four sizes · 2026-09-25
What $0.80 and 151 seconds producedthe parsed GLB
triangles493,418
vertices262,710
embedded PBR textures3
it is realabout sixteen megabytes that parse as valid glTF 2.0, UV unwrapped, with exactly three embedded PBR textures, on both subjects tried, usable as it comes rather than as a starting point
the clock is shortnine hundred seconds on that link, confirmed in the payload, the tightest expiry in either aisle, so download inside the same run
all three formats come in one joban earlier version of this page said asking for OBJ or STL returns a 404, that no converter exists and that another format is another $0.80. All three are wrong: the outputs parameter takes glb, obj and stl together, the schema says outright that OBJ and STL are converted from the GLB, and one $0.80 run returned all three
and the expiry is not a deadlinethe download endpoint re-mints a fresh fifteen-minute link for $0.00, so a missed window costs a round trip rather than the job
the documented default is rejectedthe face-count parameter prints a default in its own schema that the provider refuses; the values it accepts are a short list well above it, so following inspect literally returns a 400
and it takes time151.4 seconds, which is long enough that a job runner needs to expect it rather than time out
$0.80one mesh, 151 seconds
A hundred images
Say this to your agent
>work out what a hundred images costs on each route at the tier i actually need, not at the cheapest tier. the ranking between vendors reverses between the low and the high setting
What the agent runsdone
alibaba /qwen-image$0.030
alibaba /qwen-image-plus · not run$0.030What came backone prompt, four sizes · 2026-09-25
A hundred imagesUS dollars
cheapest vendor$0.35
flat-price vendor, any size$3.00
top tier on the other$7.50
the range is wider than this page first measureda hundred images ran $0.35 at the cheapest vendor and about $21 at the dearest tier, against the $0.60 to $7.50 an earlier version quoted. The point is the span, not the endpoints: price the tier you actually need
where it crossesat the MIDDLE quality setting of the token-billed vendor. Below it the token-billed one is 5.0x cheaper; at medium the flat-price one is 1.76x cheaper; at high the token-billed one is about 7x dearer
the flat-rate bargainone vendor charges the same $0.030 whether you ask for 1024 square or 2048 square, so at large sizes it is the cheap option rather than the dear one
count the images you getasking for two returns one entry with both images inside a content array, so iterating the entries drops one you paid for
$0.00arithmetic on measured rates
What these two aisles cannot do
Say this to your agent
>before you design a pipeline, tell me what is genuinely missing and what merely needs a step: vector output and rigging are absent, transparency is a parameter on the image endpoints rather than a separate remover, and a generated image reaches the mesh generator by downloading it and putting it through the upload endpoint. check each one against the schema rather than against a capability search
What the agent runsdone
suzanne /photo-to-3dcatalog rate
suzanne /text-to-3d$0.80What came backone prompt, four sizes · 2026-09-25
the two aisles DO connect, for $0.015an earlier version of this page said there is no path from a generated image to a mesh. There is: an upload endpoint takes the file and hands back the id the photo-to-3D call wants, confirmed end to end with a generated PNG. The honest sentence is download and re-upload for a cent and a half, not that it cannot be done
and transparency is a parameterevery endpoint on one vendor takes a transparent background directly, so a separate background remover is not something this aisle is missing
what is genuinely absentvector output and rigging. There is no upscaler either, though one endpoint re-renders to 4K, which covers some of what a reader wants an upscaler for
and it is not PNG onlythirteen image endpoints rather than eleven, and one returns JPEG over a plain http url
what discover actually doesit scores about 0.79 for upscale and for remove background, and on a second look both of those hits can do the job asked of them, so the warning is to open the schema rather than that the results are junk
$0.00know the holes first
Read more
View all →How do I give an agent a voice?speak it, read it back, and check the transcript
Step 1
Set up Monid.
One line. It installs the CLI and asks for an API key from app.monid.ai. New accounts start with $1.00.
Say this to your agent
>set up https://monid.ai/SKILL.md
Step 2
Pick what you need. Say it in one sentence.
Click a job. Paste the sentence, fill in the brackets.
Turn text into speech
Say this to your agent
>say <text> out loud with elevenlabs on monid. it returns a hosted link rather than audio bytes, so download the file in the same run: the link is documented to last an hour and the file a week
What the agent runsdone
elevenlabs /v1/text-to-speech$0.05 / 1k chars
elevenlabs /v1/text-to-dialogue$0.10 / 1k chars
elevenlabs /v1/voicesread-only
minimax /v1/t2a_v2catalog rateWhat came backone sentence, round trip · 2026-09-25
The same sentence, two modelsUS dollars for 69 characters
the default model$0.0069
the faster model$0.00345
what comes backa download link on Monid's own file service, never inline bytes, with a documented one-hour link expiry and a seven-day file expiry
asking again is safefetching the run a second time mints a fresh link and the old one keeps working, so a link that has expired costs nothing to replace
but the file clock does not restartthe seven days run from the original call, not from the refetch, so a fresh link on day seven buys you hours rather than another week
and there IS a pollan earlier version of this page said there is no polling and no job id anywhere here, which contradicts the platform's own documented run-then-poll loop; fetching the run by id IS that poll
one field halves it, and you must pass itnaming the faster model cut the bill exactly in half on the same words on both runs, and it produced slightly more audio, so the per-minute gap is a little wider than 2x. Left unset, the default is the dearer model, so the saving is something you ask for rather than something you get
the default voice is not in the listthe documented default voice id does not appear among the voices the voice endpoint returns, so pick one from the list rather than relying on the default. The list is account-scoped and came back with 42 and then 44 entries, so read it rather than quoting a count
$0.006969 characters spoken
Read it back and check it
Say this to your agent
>transcribe the audio you just made with elevenlabs on monid, then diff the transcript against the text you sent character by character. expect a word with a diacritic or an unusual spelling to come back wrong and expect the final full stop to be missing. try that word in keyterms, but check the result: on our second run keyterms turned a wrong word into a non-word rather than fixing it
What the agent runsdone
elevenlabs /v1/speech-to-text$0.22 / hour
elevenlabs /v1/forced-alignment$0.22 / hour
elevenlabs /v1/text-to-speech$0.05 / 1k charsWhat came backone sentence, round trip · 2026-09-25
it came back wrongon a clean machine-generated clip with no noise, no accent and no crosstalk, a word came back as a different word and the closing full stop was dropped; that dropped full stop reproduced on a second sentence in all three runs
what breaks is not always the brand nameon a second sentence the brand name and the unusual surname both survived and the word with a diacritic did not: Wroclaw came back as Rockwool
and keyterms does not reliably fix itthe first run it did. On the second it made it worse, turning the wrong word into a non-word, at the same $0.000069 extra. So treat keyterms as worth a try and then check, not as the fix
what that impliesif a synthetic clip mis-transcribes a proper noun, a real recording will too; tell the transcriber the names it should expect rather than checking afterwards
the round trip is cheaptext out loud and back into text cost $0.003756 in total, which is the cheapest verification step on this page
$0.003756the whole loop
Making audio against reading it
Say this to your agent
>before you plan a job, measure what a minute of audio costs to make and what a minute costs to read, on the models you will actually use: characters per minute is not a constant, and the model this command picks by default is the expensive one. then multiply both out to a thousand minutes
What the agent runsdone
elevenlabs /v1/text-to-speech$0.05 / 1k chars
elevenlabs /v1/speech-to-text$0.22 / hour
elevenlabs /v1/music$0.15 / minuteWhat came backone sentence, round trip · 2026-09-25
One minute of audioUS dollars a minute
compose music$0.15
speak it$0.0469
transcribe it$0.003667
the multiplereading is roughly twelve to forty times cheaper than making, depending which model and which route; on defaults it came out at about 27x
the rate is not a constant, so measure itcharacters per minute of delivered audio came out at 938 on the first sentence and at 869 and 976 on a second, for the two models, so the same text buys a different number of minutes depending on the voice
and the default model is the dear onethe halving only happens if you pass the model explicitly; left to the default, a thousand minutes of speech is about $97 rather than about $43
music has two shapes, not oneone provider bills per minute and the other bills a flat price per song regardless of length, which came out at about $57 a thousand minutes against about $150, so the cheaper composer depends entirely on how long your pieces are
per character ranks differentlythe two-voice route costs the same per character as the single-voice one and less per minute, because two speakers with a laugh in them talk more slowly
$0.00arithmetic on measured rates
Music and a sound bed
Say this to your agent
>make me <n> seconds of <mood> music with elevenlabs on monid, and a <n> second sound bed. this provider bills exactly the duration you ask for, so ask for exactly what you need. if you use the other music provider instead, note that it takes no duration parameter at all and bills a flat price per song, so short pieces are dear there and long ones are cheap
What the agent runsdone
elevenlabs /v1/music$0.15 / minute
elevenlabs /v1/sound-generation$0.12 / minute
minimax /v1/music_generationcatalog rateWhat came backone sentence, round trip · 2026-09-25
it is deterministicyou set the duration and it bills exactly that duration: $0.025 for ten seconds of music and $0.01 for a five-second bed, with no rounding surprise
so budget in secondsunlike the character-priced routes, there is nothing to measure afterwards; the number you ask for is the number you pay
the second vendor works and prices differentlyan earlier version of this page wrote it off as dead. On a second run it answered and billed a flat fifteen cents for a song of about two and a half minutes, with no duration parameter to set
a third one is the dead onea different provider's music endpoint, flagged unknown in the catalog rather than stable, timed out with a 504; it billed nothing, as every genuine failure here did
so ask for exactly what you need is provider-specificit is true of the per-minute provider and meaningless on the flat-rate one
$0.035ten seconds and five
One file, four different durations
Say this to your agent
>if you send the same audio to more than one endpoint, record what duration each one billed: on our runs one identical file was billed at four different durations by four endpoints on the same provider, spreading about fifty percent. one of them also enforces a 4.6 second minimum that is in no schema, and the CLI hides that rejection behind a generic 400 with no output file, so read the run record rather than the client error
What the agent runsdone
elevenlabs /v1/speech-to-text$0.22 / hour
elevenlabs /v1/forced-alignment$0.22 / hour
elevenlabs /v1/speech-to-speech$0.12 / minute
elevenlabs /v1/audio-isolation$0.12 / minuteWhat came backone sentence, round trip · 2026-09-25
The same 4.133 second file, billedseconds charged
transcription6 seconds
forced alignment5 seconds
isolation4 seconds
voice conversion4 seconds
the spreadone file, four endpoints on the same provider, four different billed durations: 67% apart on the first file and about 50% on a second, so the shape holds and the percentage does not
and the floor contradicts itselfthe endpoint that refuses a 4.55 second file for being under its 4.6 second minimum billed a 5.108 second file as four seconds, which is below the minimum it just enforced
why it matters at scalea pipeline that estimates cost from the file's real duration will be wrong in three different directions depending on which endpoint it used
the undocumented floorone endpoint enforces a 4.6-second minimum that appears nowhere in its schema, and rejects anything shorter
and the CLI hid itthat rejection printed only a generic HTTP 400 and wrote no output file, while still creating three real runs whose records carried the actual too-short message
$0.00the spread, not the cost
What this aisle cannot do
Say this to your agent
>tell me plainly what is not here before i design around it: there is no dubbing, and nothing in these three aisles streams a generated file back to me. do NOT conclude from that that live voice is impossible: a live phone call is metered by the second elsewhere on the platform, and the run stays open for the length of the call. check whether a cloned voice can be USED before you tell me one cannot be MADE: they are different questions
What the agent runsdone
elevenlabs /v1/voicesread-only
elevenlabs /v1/speech-to-speech$0.12 / minute
minimax /v1/t2a_v2catalog rateWhat came backone sentence, round trip · 2026-09-25
no voice CREATION, which is not the same as no cloned voicesnothing in these aisles creates or clones a voice. But one provider's own voice setting states that system, cloned and AI-designed voices are all accepted, and it carries a timbre control that blends four voices into a new one, so a cloned voice made elsewhere can be used here
and the voice list is not a fixed 42it returned 44 on a second run and it is scoped to the account, so it is a list to read rather than a number to quote
no streaming, and that does NOT rule out live voiceno endpoint in these three aisles streams a generated file back to you; every route is request, wait, download. An earlier version of this page concluded from that that a real-time conversational agent is impossible. It is not: a phone call elsewhere on the platform is metered per second at about thirty cents a minute, the run stays open for the length of the call and stopping the run hangs up
a stream flag exists but is pinned offtwo endpoints do declare a stream parameter, fixed at false, so an agent grepping the schemas will find one and must read the constant rather than the key
no dubbingthere is a voice-to-voice conversion endpoint, and it is not the same thing as translating and re-voicing a track
read the notes, not the summaryfive of nine summaries describe the output as base64 audio and all five returned a hosted link instead; the notes block was the accurate one, and it documents base64 as a fallback for when the file cannot be saved rather than as something that never happens
$0.00know the holes first
