How-to · 3 steps

How to pick a model
with Monid.

Give this to your agent
$set up https://monid.ai/SKILL.md, then tell me what the boards say about <model> and what it costs per task

and let it take it from there.

›which model should i run this on
#JobWhat came backCost
1The code board134 models, crowd Elo
2The chat boarda different model at #1
3157 models pricedcost and an index each
4One model, 5 hosts3x on price, 2.2x on speed
5The empty one200 OK, just a slug
6And againsame empty, second slug
7Judge five rowsone call, 1,342 tokens
8Embed six strings128 tokens, real vectors
9The 502twice, billed nothing
10Totalfifteen runs, thirteen billed
two boards · one price table · two models at number oneagent running$0.1301

Real pull, 2026-09-25. Each board named, each figure dated.

Step 1

Set up Monid.

One line. It installs the CLI and asks for an API key from app.monid.ai. New accounts start with $1.00.

Say this to your agent
>set up https://monid.ai/SKILL.md
Step 2

Pick what you need. Say it in one sentence.

Click a job. Paste the sentence, fill in the brackets.

What the coding board says today

Say this to your agent
>get the lmarena code leaderboard from monid and give me the top ten with their ratings and vote counts. label it as crowd elo on coding tasks and tell me today's date, because the payload has no timestamp in it
What the agent runsdone
lmarena /get_code_leaderboard$0.01
lmarena /get_best_coding_models$0.01
lmarena /get_top_models_by_category · not run$0.01
lmarena /get_vision_leaderboard · not run$0.01
lmarena /get_text_to_video_leaderboard · not run$0.01
What came backone pull · 2026-09-25
Code Arena rating, pulled 2026-09-25crowd Elo on coding tasks
1st · claude-opus-5.5-max1826.7
2nd · gpt-6-astra-max1791.7
134th · devstral-medium1075.4
what this iscrowd Elo from head-to-head votes on coding prompts, run by LMArena, as it stood on 2026-09-25
what it is nota measurement of your task; the prompts are theirs, the voters are theirs, and neither is your codebase
the spread751 rating points between first and last across 134 models, which is a real gap, unlike the gaps at the top
the missing fieldno timestamp anywhere in the payload, so the pull date has to be recorded by you, not read off the response
$0.01134 models, one call

When two boards crown two models

Say this to your agent
>search the lmarena text leaderboard for <lab> on monid, passing category explicitly, and give me every entry with its rating and its rank interval. if the intervals overlap, say so instead of calling one of them better
What the agent runsdone
lmarena /search_leaderboard_models$0.01
lmarena /get_best_coding_models$0.01
lmarena /get_text_leaderboard · not run$0.10
lmarena /get_leaderboard_overview · not run$0.10
What came backone pull · 2026-09-25
Text Arena, top two, same daycrowd Elo on chat
claude-fable-5-high1505.68
claude-opus-4-6-high1504.56
the disagreementthe same provider's code board and chat board, pulled the same day, put two different models from the same lab at number one
the gap at the top1.12 rating points between first and second, with the rank intervals of the top seven all reading 1 to 7, so the order is not distinguishable
the $0.10 controlwe paid for the full table to check the cheap search: 402 entries, 31 from that lab, exactly the same model keys the $0.01 call returned. Ten times cheaper, nothing lost
pass category anywaythe field has no enum and silently defaults to text, so asking for the code arena and forgetting it returns chat numbers labelled as code. A typo returns an empty success at full price
$0.01seven rows from one lab

What a task actually costs, across every model

Say this to your agent
>get cost per task from artificial analysis on monid for the whole default set, then show me cost against the intelligence index and drop any row whose cost breakdown is all zeros
What the agent runsdone
artificial-analysis /get_cost_per_task$0.05
artificial-analysis /get_model_cost_per_task · not run$0.10
artificial-analysis /get_models_list · not run$0.02
artificial-analysis /get_model_detail$0.01
What came backone pull · 2026-09-25
Cost per task against the index, pulled 2026-09-25US dollars per task · index score
claude-fable-5$8.746 · 49.6
claude-opus-5-5$5.982 · 57.6
gpt-6-astra$3.258 · 52.7
gpt-6-astra-xhigh$2.309 · 52.4
mimo-v2-6-pro$0.133 · 46.3
gpt-6-luna-low$0.0045 · 20.9
the crossingthe most expensive row in the 157 is not the highest-scoring one; a model at two thirds the cost scores eight index points above it
the cheap endone model costs $0.0045 a task at an index of 20.9, which is the best score per dollar in the set and nowhere near the best score
the endpoint that took the moneythe per-model detail call billed $0.01 and returned only the slug, twice, on two different models
never sum the breakdownits five components add up 67% above the total they belong to, because a cache hit in that list is a saving and not a charge; quote cost_per_task, never a reconstruction
$0.05157 models, one call

Same weights, five different bills

Say this to your agent
>get the providers for <open model> from artificial analysis on monid and give me price per million in and out, median output speed and median time to first token for each host. flag any host that is missing a performance field instead of treating it as zero
What the agent runsdone
artificial-analysis /get_model_providers$0.01
artificial-analysis /search_models$0.01
artificial-analysis /get_llm_providers_leaderboard$0.01
artificial-analysis /get_hardware_benchmarks · not run$0.03
What came backone pull · 2026-09-25
One open-weight model, five hostsUS dollars per 1M input tokens
GMI · speed not reported$0.348
Xiaomi · 46.9 tok/s$0.435
Novita · 43.5 tok/s$0.522
DeepInfra · 37.3 tok/s$1.000
Zyphra · 80.5 tok/s$1.000
the spread3x on input price and 2.2x on output speed for identical weights, all from one $0.01 call
cheapest is not fastesttwo hosts charge exactly the same and one is more than twice as fast, with a time to first token 2.3x lower
the missing fieldsthe cheapest host reports real prices and no speed numbers at all, so it can be priced and not timed from this response
the failurethe aisle's provider leaderboard 502'd twice and billed $0.00 both times, so this call is the route that works today
$0.01five hosts, one call

Score your own rows instead of trusting a board

Say this to your agent
>judge these rows with typesafe on monid, all of them in one call, and send the model field explicitly even though the schema says it has a default. give me the score and the probability distribution per row
What the agent runsdone
typesafe /systemone$0.042 / 1M input
artificial-analysis /get_cost_per_task$0.05
What came backone pull · 2026-09-25
A five-row judge batchscore out of 3, against hand-written quality
row 1 · deliberately poor0.26
row 21.86
row 31.92
row 42.95
row 5 · deliberately strong2.99
the price1,342 input tokens for the whole batch, 56 micro-dollars, which is $0.0000112 a row
why batching winsonly input tokens bill and the shared state is ingested once, so the per-row cost falls as the batch grows
what it showedthe scores rose in step with the quality we wrote into the rows, which is the check worth running before you trust any judge
two schema trapsthe model field is documented with a default and omitting it 422s every time; and the scale is zero-indexed, so a four-level rubric scores 0 to 3 and a 2.99 is nearly perfect, not a 75%
$0.000056five rows, one call

Embed a corpus for a hundredth of a cent

Say this to your agent
>embed these strings with octen on monid with model set to octen-embedding-0.6b, and give me the input token count. work the cost out from those tokens and the rate, because the run reports $0.00 for amounts this small
What the agent runsdone
octen /embedding$0.01 to $0.07 / 1M
typesafe /systemone$0.042 / 1M input
What came backone pull · 2026-09-25
name the size exactlythe field defaults to the 4b model at four times the rate, and "the 0.6b size" is not a value the API takes; the literal enum is octen-embedding-0.6b and nothing warns you
per item$0.000000167 a string at the small size, which is a sixth of a millionth of a dollar
at the top sizethe 8b tier is $0.07 a million tokens, so the same 128 tokens would be $0.000009, still under a thousandth of a cent an item
where the cost comes fromthe run's own cost field rounds these to $0.00, so the only honest number is the billed input tokens times the rate; count tokens, not calls
$0.000001six strings, 128 tokens
Step 3

Take the cheapest route with Monid.

Every endpoint in this aisle bills per call, so a miss costs what a hit costs. Ask in the order that answers your question soonest.

Resolve the slug · $0.01
Search models$0.01find the slug
Search the boards$0.01find it on a board
What the boards say · cheapest first
Code Arena$0.01
Models list$0.02
Hardware benchmarks$0.03
Cost per task$0.05
Text Arena, full$0.1
Who hosts it
Model providers$0.01
Judge your rows$0.0000112 / rowwhen the board is not your task
Back to you
rowanswered bycost
1Code Arena$0.03
2Cost per task$0.13
3not tracked$0.01
4Code Arena$0.03
5Models list$0.05
5 of 5 rows$0.25
Say this to your agent
>tell me what the boards and the price table say about these models, cheapest call first, and show the cost per model: <paste model names>
“cheapest call first”every endpoint here bills per call, so a wrong guess costs the same as a right one
“name the board”the code board and the chat board crowned two different models on the same day
“read the rank interval”the top seven of one board all report a rank of 1 to 7, which means tied
“drop the all-zero rows”four of 157 carry a real index next to a cost breakdown of nothing but zeros

One key. 1,700+ tools.

Three small aisles: 22 benchmark endpoints, one judge, one embedder.

Code Arena, crowd Elolmarena · $0.01 / call
Text Arena, the full tablelmarena · $0.10 / call
Vision, image and video boardslmarena · $0.01 / call
Find one model in a boardlmarena · $0.01 / call
157 models, cost and indexartificial analysis · $0.05
Who hosts it, for how muchartificial analysis · $0.01
Every model they trackartificial analysis · $0.02
The silicon underneathartificial analysis · $0.03
An LLM judge, per tokentypesafe · $0.042 / 1M input
Embeddings in three sizesocten · $0.01 to $0.07 / 1M

+ 1,700 more across 55 providers

Browse the catalog →