How to Use Seedance 2.0 by API: One Prompt, Two Models, Measured
Seedance 2.0 and 2.0 Mini rendered one 5-second prompt in the same two minutes for the same 108,900 tokens. One billed exactly twice the other.

Copy this line to your agent to generate a video with Seedance 2.0.
set up https://monid.ai/SKILL.md and use bytedance /v1/video/seedance-2.0-mini for a text prompt
On 2026-09-15 we sent one prompt to Seedance 2.0 and to Seedance 2.0 Mini, at the same resolution, the same length and the same aspect ratio. Both took about two minutes. Both billed 108,900 tokens. One cost exactly twice the other, and nothing in the response tells you what the second half of that price bought. This guide runs through Monid, the OpenRouter for agent tools.
What is Seedance 2.0 and which version should I use?
Seedance 2.0 is ByteDance's text-to-video and image-to-video model, and the version question has a shorter answer than the pricing page suggests: start on Mini, and move to 2.0 when you can see why.
The two endpoints
Both are the same model family behind the same request shape:
bytedance /v1/video/seedance-2.0 480p, 720p, 1080p, 4k
bytedance /v1/video/seedance-2.0-mini 480p, 720p
The body is identical. A content array holding text or image references, a resolution, a duration between 4 and 15 seconds, an aspect ratio, and three switches: generate_audio, watermark, priority.
What "Mini" does not mean
It does not mean faster. Our two runs completed in 125 and 124 seconds. It does not mean fewer tokens. Both billed 108,900 for a 5-second 720p clip. It means a lower per-token rate and a ceiling at 720p, and a rendering that is visibly softer in a way this article shows rather than asserts.
The rule
If 720p is your delivery resolution, render on Mini first. If the result is good enough, you are done at half the rate. If it is not, or you need 1080p or 4k, that is what 2.0 is for. Rendering on 2.0 by default because it is the "real" one is paying double for a difference you have not looked at.
📖 See also MiniMax vs Seedance for AI Video
How do you use Seedance 2.0 through an API?
Three steps, and the first two cost nothing.
For agents
Grab an API key at app.monid.ai, then paste this to your agent and hand it the key:
set up https://monid.ai/SKILL.md
It learns the whole discover, inspect, run workflow itself. More in the agent quickstart.
For humans
npm install -g @monid-ai/cli
monid keys add -k <your-api-key> -l main
Step 1. Read the schema and the price matrix
What it does. Shows you the exact body shape and how the bill is computed before you spend anything.
The call.
monid inspect -p bytedance -e /v1/video/seedance-2.0-mini
What comes back. The content array shape, the resolution enum, the duration bounds, and a pricing matrix keyed on resolution with a note on how tokens are estimated. Read the note. It is the only place the token formula is written down.
What it costs. Nothing. Inspect is free on every endpoint.
Step 2. Generate on Mini
What it does. Renders the clip and returns a download link.
The endpoints. bytedance/v1/video/seedance-2.0-mini, billed per token, resolution is the price selector.
The call.
monid run -p bytedance -e /v1/video/seedance-2.0-mini -w \
-i '{"content":[{"type":"text","text":"A small paper boat drifting across a puddle on a city sidewalk after rain, low camera angle, soft morning light, gentle ripples, cinematic"}],"resolution":"720p","duration":5,"ratio":"16:9","generate_audio":false}'
-w waits inline. Generation takes tens of seconds to minutes; ours took 125.
What comes back. A video_url, a model id, and a status of succeeded. The run record carries billedUnits, which is the token count you were charged for.
What it costs. A fraction of a dollar for a 5-second 720p clip, billed on actual tokens. Current figures at monid.ai/tools.
Step 3. Download inside 24 hours
What it does. Keeps the file.
The call.
curl -L "<video_url>" -o clip.mp4
ffprobe clip.mp4
What comes back. Ours: 1280 by 720, 24 frames per second, 5.04 seconds, 2.7 MB. The URL is signed and expires in roughly a day, which the endpoint notes say plainly and which a pipeline that stores the URL instead of the file will discover the hard way. This is the same download-promptly discipline as any signed asset URL, and the reason the short-form video pipeline writes the file to storage as its first step after generation.
Give this to your agent![]()
Set up https://monid.ai/SKILL.md, and then use Monid to render this prompt on seedance 2.0 mini at 720p for 5 seconds, download the result, and tell me how many tokens it billed.
What does a Seedance run actually cost and how is it billed?
Per token, where a token is a unit of pixels over time, and the formula is in the endpoint notes: tokens are roughly width times height times 24 frames times seconds, divided by 1,024.
The measurement
Two runs, 2026-09-15, identical body except the endpoint:
| Seedance 2.0 Mini | Seedance 2.0 | |
|---|---|---|
| Resolution, duration, ratio | 720p, 5 s, 16:9 | 720p, 5 s, 16:9 |
billedUnits (tokens) | 108,900 | 108,900 |
| Wall clock | 125 s | 124 s |
| Output | 1280x720, 24 fps, 5.04 s | 1280x720, 24 fps, 5.04 s |
| Bill relative to Mini | 1.0x | 2.0x |
The formula predicts 720p at 16:9 costs about 21,600 tokens a second, so 5 seconds is about 108,000. We were billed 108,900. The estimate is good to within one percent.
What that means for budgeting
Duration is linear. A 10-second clip is twice the tokens of a 5-second clip. No surprise, but it means duration is your first cost control and the 4-to-15 range spans nearly a 4x bill.
Resolution is not linear in the way you expect. The per-token rate in the matrix is the same at 480p and 720p on 2.0, slightly higher at 1080p, and lower at 4k. That last one reads like a discount and is not: 4k has roughly nine times the pixels of 720p, so roughly nine times the tokens, and a lower rate on nine times the count is a much larger bill. Read the token column, not the rate column.
Audio is a switch. generate_audio defaults to true. We turned it off for the comparison. Leave it on only when you want it, because it is part of what the model does and it is not free to do.
Why the two bills differ by exactly 2x
Because Mini's per-token rate at 720p is half of 2.0's, and the token count was identical. There is no hidden efficiency: the endpoints do not tokenise the same clip differently. The whole price gap is the rate, and the rate is buying rendering quality, which is the next section.
What is the difference between Seedance 2.0 and 2.0 Mini?
Rendering quality and the top two resolutions. Not speed, not token efficiency, not the request shape.
What the API exposes
Two things only. 2.0 offers 1080p and 4k; Mini stops at 720p. And 2.0's rate is higher. Everything else in the schema is identical, and the run records for our two clips are identical in every field except model id, endpoint and cost.
What you can only see by looking
We pulled the frame at 2.5 seconds from each clip.
The Mini render is warm, with shallow depth of field and heavy bokeh in the background, the puddle edged with wet cobblestones and the boat sitting at the water's edge. It reads as a lens shot with a wide aperture.
The 2.0 render is cooler and wider, the pavement texture sharper across the whole frame, the ripple rings around the boat cleaner and more concentric, and the boat itself more literally a folded paper boat. It reads as a more controlled composition.

Both honoured every element of the prompt. Neither is wrong. Whether the second is worth twice the first depends entirely on where the clip is going, and that is a judgement the API cannot make for you, which is why the rule above is to render Mini first and look.
The honest limit of this measurement
One prompt, one seed, one frame each. Generative models vary run to run, and a different prompt could narrow or widen the gap. What does not vary is the token count and the rate, and those are the part of the comparison you can budget on. The rendering difference is real and it is one sample.
Which endpoint should I use for which job?
| Endpoint | What it does | Input | Output | Best for | Billing |
|---|---|---|---|---|---|
bytedance/v1/video/seedance-2.0-mini | Text or image to video, up to 720p | content[], resolution, duration, ratio | Signed video_url, 24 h | The first render of anything | Per token, by resolution |
bytedance/v1/video/seedance-2.0 | Same, up to 4k | Same body | Same shape | 1080p and above, or when Mini's look is not enough | Per token, higher rate |
bytedance/v1/video/seedance-2.0-fast | Lower-latency variant | Same body | Same shape | Interactive use where the two-minute wait matters | Per token |
bytedance/v1/video/seedance-2.5 | The newer model | Same body | Same shape | When 2.0's output is the ceiling | Per token, highest rate |
lmarena/get_text_to_video_leaderboard | Where these sit against other models | none | Ranked list | Deciding whether to use Seedance at all | Per call |
Every row was verified with monid inspect on 2026-09-15. The table gives billing shape rather than figures, because shape drives design and current numbers live on monid.ai/tools.
The last row is there because "which video model" is a question with a public answer that updates, and a leaderboard call costs less than one wrong render. The voice side of the same generative media catalog is compared in ElevenLabs alternatives.
When is Seedance the wrong tool?
Four cases.
You need a specific real person on screen. Reference images containing real human faces are restricted by the provider, and the model is not built to hold a consistent identity across clips. For a fixed presenter or actor you want a different class of tool, and we have written up what these video endpoints cannot do alongside what they can.
You need it in under a minute. Two minutes for five seconds of video is the measured latency. The -fast variant exists for a reason, and even that is not interactive. Anything that blocks a user on generation needs a queue and a callback, not a wait.
You are storing the URL. It expires in about a day. If your system keeps the link rather than the bytes, every clip older than that is gone. This is not a Seedance quirk; every signed-URL media API works this way, and the fix is one download step.
Your budget is per clip, not per second. Billing is per token, tokens scale with pixels and seconds, and a 15-second 4k render on 2.0 is a very different line item from a 5-second 720p render on Mini. Estimate with the formula before you loop.
And the disclosure: this is Monid's blog and we resell both endpoints. The recommendation here is to use the cheaper one by default and pay for the expensive one only after you have looked, which is a recommendation to spend less.
Conclusion
How to use Seedance 2.0 is a two-line request body and a two-minute wait, and the question underneath it is which of the two endpoints to send that body to. On our measurement the answer is Mini first, because it produced the same clip length at the same resolution in the same time for the same token count at half the rate, and the difference 2.0 buys is one you have to watch to see.
What matters more than the model choice is understanding the bill. Tokens are pixels over time, the formula is in the schema notes and it was accurate to within one percent, duration is linear, and the 4k rate that looks cheaper is attached to nine times as many tokens. A pipeline that estimates with the formula before it renders will never be surprised by a video bill.
Free next step: run monid inspect -p bytedance -e /v1/video/seedance-2.0-mini and read the pricing note, then render one 5-second clip and compare billedUnits to the formula. Start at monid.ai.
FAQ
Where can I use Seedance 2.0?
Through ByteDance's own platforms, through third-party video tools that have integrated the model, or by API through a provider that wraps it, which is what this guide covers. The API route is the one that fits inside a pipeline: one request body, one signed URL back, billed per token on one balance alongside every other tool your agent calls. Discovery and inspect are free, so checking whether the endpoint does what you need costs nothing before the first render.
Can Seedance 2.0 do image to video?
Yes. The content array accepts image_url items alongside or instead of text, for first-frame or reference-image generation. Two constraints from the endpoint notes: reference URLs must be public https links or an asset:// id, not inline base64, and reference images containing real human faces are restricted. For product shots, scenes and objects it works as documented; for a consistent human presenter it is the wrong tool.
Does Seedance generate audio?
It can, and it does by default: generate_audio is true unless you set it false, and it produces synchronised voice, sound effects and music. We turned it off for the cost comparison so the two runs were measuring video alone. If you are layering your own voiceover or music afterwards, turn it off; if you want the clip to arrive finished, leave it on and expect it to be part of the bill.
Why did my video link stop working?
Because the video_url is a signed link that expires in roughly 24 hours, which the endpoint notes state directly. The clip is not gone from the provider's storage the moment the link dies, but you have no way to ask for it again through the same run. Download the file as the first step after a run completes and store the bytes; treat the URL as a delivery mechanism, not a reference.
Last updated September 2026.


