LLM as a Judge, or a Model That Only Judges: Jev Measured
Four typed judgments on one review in 1.9 seconds, billed on 631 input tokens with output free. The severity score came back 1.1, and it is zero-indexed.

Copy this line to your agent to get typed judgments your code can branch on.
set up https://monid.ai/SKILL.md and use typesafe /systemone with noul, choice and score questions
We asked one model four questions about a one-star product review on 2026-09-24. It came back in 1.9 seconds with a probability, a classification with its full distribution, a severity value of 3 out of levels numbered 0 to 4, and a bill computed from 631 input tokens with the 120 output tokens free. Then we handed it a deliberately ambiguous review and the same severity question returned 1.1. A model that only judges behaves differently from a model you ask to judge, and that difference is measurable. This guide runs through Monid, the OpenRouter for agent tools.
What does LLM as a judge actually mean?
Two quite different things, and the phrase covers both.
The pattern
Using a language model to evaluate something rather than to write something: score this answer, classify this ticket, decide whether these two records are the same person. It became standard because evaluation is the step that scales badly with humans.
The usual implementation
Prompt a general model, ask for JSON, parse the JSON, hope. It works, and it carries three taxes: you are paying generation prices for a decision, you are parsing free text that occasionally is not valid JSON, and the number it gives you is a token the model chose rather than a calibrated quantity.
The other implementation
A model built only to answer typed questions. You hand it a state and a map of questions, each declared as a type, and it returns values in those types with probabilities attached. No prompt engineering, no parsing, and the output shape is a contract rather than a hope.
typesafe/systemone is the second kind, and it appeared on our shelf recently enough that our own research note from four days earlier said it was not there yet. Three question types: a yes/no with a probability, a classification into your own options with a full distribution, and an ordered score with per-level probabilities.
Why the distinction is worth a measurement
Because "it returns JSON" is not the interesting part. The interesting part is whether the numbers mean anything, and the next section is two reviews that answer that.
📖 See also Are Instagram Follower and Engagement Trackers Accurate?
What does a model that only judges return?
Typed answers with their uncertainty exposed, and the uncertainty moves when the case does.
The clear case
A one-star review describing a motor that started grinding after three weeks. Four questions, one call, 1,853 milliseconds:
{ "broke_in_use": { "type": "noul", "noul": 0.93 },
"failure_mode": { "type": "choice", "choice": "mechanical_failure",
"confidence": 1,
"probabilities": { "mechanical_failure": 1, "underperformed": 0,
"arrived_damaged": 0, "wrong_expectation": 0,
"packaging": 0 } },
"severity": { "type": "score", "score": 3, "confidence": 0.98,
"probabilities": { "2": 0.01, "3": 0.98, "4": 0.01 } },
"mentions_value": { "type": "noul", "noul": 0.94 } }
The distribution on the classification collapsed completely: probability 1.0 on one option, zero on the other four, confidence exactly 1. That is a model saying there is nothing to weigh here.
The unclear case
A second review, same product, where the buyer says it works but is not as good as the videos and wonders whether they are using it wrong. Same four questions, 2,083 milliseconds:
{ "broke_in_use": { "noul": 0.06 },
"failure_mode": { "choice": "wrong_expectation", "confidence": 0.82,
"probabilities": { "wrong_expectation": 0.86,
"underperformed": 0.14 } },
"severity": { "score": 1.1, "confidence": 0.71,
"probabilities": { "0": 0.12, "1": 0.65, "2": 0.23 } },
"user_error": { "noul": 0.97 } }
broke_in_use went from 0.93 to 0.06 on the same question. Confidence on the classification dropped from 1 to 0.82, with 14 percent still sitting on the second-best option. Severity dropped from 0.98 confidence to 0.71 with the mass spread across three levels.
That is what calibration looks like when it is real. The model did not just change its answer, it changed how sure it was, and it showed you where the remaining doubt sits.
The field that is the bill
"usage": { "input_tokens": 631, "output_tokens": 120 }
Billing is on input tokens only. Output is free. Four judgments on one review cost 631 tokens of input, which at this endpoint's rate is a small fraction of a cent, and the second review cost 583.
How do you run typed judgments through one key?
One call, three question shapes, and one parameter mistake worth skipping.
For agents
Grab an API key at app.monid.ai, then paste this to your agent and hand it the key:
set up https://monid.ai/SKILL.md
It learns the whole discover, inspect, run workflow itself. More in the agent quickstart.
For humans
npm install -g @monid-ai/cli
monid keys add -k <your-api-key> -l main
The request shape
The endpoints. typesafe/systemone, billed per input token, body takes state, model and questions.
state is what you are judging: a string, or better, an object so each part of the context has a name. questions is a map whose keys become the keys of your answer object.
{ "state": { "review_body": "…", "rating": 1, "product": "milk frother" },
"model": "jev-latest",
"questions": {
"broke_in_use": { "type": "noul",
"instructions": "Did the product physically fail during normal use?",
"criteria": { "true": "Stopped working while used as intended.",
"false": "Never worked, arrived damaged, or the complaint is expectations." } },
"failure_mode": { "type": "choice",
"instructions": "What is the primary failure mode?",
"criteria": { "mechanical_failure": "A part wore out or seized.",
"arrived_damaged": "Damage before first use.",
"underperformed": "Works but not well enough." } },
"severity": { "type": "score",
"instructions": "How severe is this for the customer?",
"criteria": ["trivial annoyance", "minor inconvenience",
"noticeable problem", "product unusable", "unsafe or total loss"] } } }
The two field names that cost us a request
Our first attempt used question and options. The real fields are instructions and criteria, and criteria changes shape by type: an object keyed by your option names for choice, an array of two to ten ordered levels for score, and an optional true/false pair for noul. The 400 listed every wrong key and charged nothing, which is the good kind of failure.
Batch, because the state is the cost
The endpoint's own pricing note says the state is ingested once per call. Four questions on a 631-token state cost 631 tokens. Asking those four questions in four calls costs roughly four times that for the same answers. One call per item, many questions per call.
Give this to your agent![]()
Set up https://monid.ai/SKILL.md, and then use Monid to for each of these 500 one-star reviews, decide whether the product broke in use, classify the failure mode into five categories, score the severity, and give me the three most common failure modes with counts.📖 See also The Best Amazon Reviews API in 2026
Why is the severity score 1.1 and not 2?
Because the scale is zero-indexed and the value is an expected value, not a level.
The legend tells you
Every score answer carries a legend mapping the numbers to the labels you supplied:
"legend": { "0": "trivial annoyance", "1": "minor inconvenience",
"2": "noticeable problem", "3": "product unusable",
"4": "unsafe or total loss" }
We passed five levels. They came back as 0 through 4. So a severity of 3 is the fourth level, "product unusable", not the third. Read the legend, never assume one-based. A dashboard labelled "severity 3 of 5" on this output is describing the wrong thing.
And the number is continuous
The ambiguous review returned 1.1. There is no level 1.1. The value is the probability-weighted average across the levels: mass of 0.12 on level 0, 0.65 on level 1 and 0.23 on level 2 works out to roughly 1.1.
That is more useful than a discrete level, and it changes how you consume it:
Rank on the score, threshold on the probabilities. A fractional score sorts a list of five thousand items sensibly. A rule like "escalate anything unusable" should read the probability mass at that level, not whether the rounded score reached it.
Never round before you aggregate. Rounding 1.1 to 1 across a large set discards exactly the signal that distinguishes a borderline population from a clear one.
Use the confidence as a routing key. Ours was 0.98 on the clear case and 0.71 on the unclear one. Sending everything under a confidence floor to a human is a cheap and honest queue design, and it is the same discipline as reading email_status before paying for a reveal in the contact pricing guide.
Which endpoint should I use for which job?
| Endpoint | What it does | Input | Output | Best for | Billing |
|---|---|---|---|---|---|
typesafe/systemone | Typed judgments on a state | state, questions, model | Probabilities, distributions, scored values with legends | Deciding, classifying, scoring at volume | Per input token, output free |
apify/delicious_zebu/amazon-product-details-scraper | The things to judge | ASIN | Product records | Supplying the state | Per result |
firecrawl/scrape | A page as text | url | Markdown plus metadata | Supplying the state from the web | Tiered per call |
hunterio/email-count | Structured company facts | domain | Counts by department | State that is already typed | Per call |
artificial-analysis/search_models | Model benchmarks | Query | Model records | Choosing a general model instead | Per call |
Every row was verified with monid inspect on 2026-09-24. The table gives billing shape rather than figures, because shape drives design and current numbers live on monid.ai/tools.
The first row judges and does not fetch. Everything above it in a pipeline is what supplies the state, which is the reason this endpoint and a tool catalog belong on the same key.
When is a general model the better judge?
Four cases.
The output is prose. If you need an explanation, a summary or a rewritten sentence, a judge model is the wrong shape. It returns values, not text, and output tokens being free is a consequence of there being very little output.
The question is open-ended. "What are the themes in these reviews" is not a typed question. Discover the categories with a general model first, then encode them as a choice and run the volume through the judge.
You need one judgment, once. The advantage here is per-item cost at scale and a parseable contract. For a single decision in a conversation, a general model you are already calling is simpler.
You cannot enumerate the options. choice takes your categories. If the category set is unknown or unbounded, you are doing extraction rather than classification, and extraction is a different job.
And the disclosure: this is Monid's blog, we resell this endpoint, and we are telling you that for prose, for exploration and for one-off decisions you should use something else. What it is genuinely good at is the case where you have thousands of things and a fixed question, and there the per-item cost and the calibrated probability are both real.
Conclusion
A model that only judges gives you two things a prompted general model does not: a contract on the output shape, and uncertainty you can read. On 2026-09-24 four typed questions on one review came back in 1.9 seconds with a probability of 0.93 on the yes/no, a classification whose distribution had collapsed to 1.0, and a severity of 3 on a scale it labelled for us. The same questions on an ambiguous review returned 0.06, a confidence of 0.82 with 14 percent still on the runner-up, and a severity of 1.1.
What matters more than the speed is reading the output correctly. The score is zero-indexed, so five levels are numbered 0 to 4 and a legend arrives with every answer. The score is also continuous, so 1.1 is a real value and rounding it before you aggregate throws away the only thing that separates a borderline set from a clear one. And the bill is usage.input_tokens alone, which means the state is what you pay for and the fifth question you add is nearly free.
Free next step: run monid inspect -p typesafe -e /systemone and read the three question types before you design a schema. Use instructions and criteria, not question and options, and you will skip the 400 we paid nothing for. Start at monid.ai.
FAQ
Why not just prompt a general model to return JSON?
You can, and for small volumes you should, since you are probably already calling one. Three things change at scale. The output is a contract here rather than a request, so there is no parse step and no occasional malformed response. The probabilities are calibrated quantities rather than tokens the model picked, which is why our yes/no moved from 0.93 to 0.06 across two reviews and the confidence moved with it. And the billing is on input only, so adding a fifth question to an existing call costs almost nothing while a fifth prompted call costs a fifth call. If any of those three does not matter to your job, prompt the model you have.
What does the probability actually mean?
For a noul it is the model's calibrated belief that the answer is yes, so 0.93 and 0.06 are meant to be read as strong yes and strong no rather than as scores. For a choice you get both a pick and the full distribution, and the distribution is the useful part: our clear case put 1.0 on one option and our unclear case left 0.14 on the runner-up, which is exactly the row a human should see. For a score you get a probability per level plus a weighted value. Treat the confidence as a routing key: everything below a floor you choose goes to a person, and the rest is automated.
How does the bill actually work?
It is input tokens only, reported in usage.input_tokens, with output tokens free. That has one important design consequence: the state is ingested once per call, so the cost is driven by how much context you send rather than by how many questions you ask about it. Our two reviews cost 631 and 583 input tokens for four judgments each. The wrong pattern is one call per question, which pays for the same state four times. The right pattern is one call per item with every question you need about that item bundled into it.
Should you pin the model version?
Pin it once your prompts and thresholds are tuned, because a judge model's calibration is the thing you built your thresholds on and a new release can shift it. The endpoint exposes three choices: a stable alias, a preview alias, and a pinned version. We requested the stable alias and the response reported back jev-1.13.0, which is the useful behaviour: the answer always names the version that produced it, so you can store it alongside every judgment and know later whether a distribution shift was your data or their release.
Last updated September 2026.


