All guides
EvaluationProduction·Advanced9 min

LLM-as-a-Judge: How Automated Prompt Scoring Works

LLM-as-a-judge means prompting one model with a rubric to score another model’s output. It is fast and repeatable enough to rank drafts at scale, but it carries documented biases toward position, length and its own writing, it reads the text it grades as input, and every score costs a model call — so validate it against human judgment before trusting it.

Scoring outputs by hand works until there are too many of them. The usual next step is to hand the rubric to a model and let it grade. This guide covers how an LLM judge works, the biases researchers have measured in judges, the practices that make one consistent, why the text it grades is an attack surface, what it costs — and, from running one in production, when a cheap heuristic is all you need.

By Andrei Bădulescu, Founder at VantagePrompt·Updated

How does LLM-as-a-judge work?

A judge is an ordinary model call with an unusual job. Its system prompt holds the rubric — what to assess and what each score means — and its input is the candidate: the output, or the prompt, being graded. It answers with a score, and usually a reason. In their MT-Bench study, Zheng et al. describe three setups:

SetupWhat the judge seesUse it for
Single-answer gradingOne candidate and a rubric; returns a score.Monitoring quality over time; scoring at scale.
Pairwise comparisonTwo candidates for the same input; picks the better one or a tie.Choosing between two prompt versions.
Reference-guided gradingOne candidate plus a reference answer to grade against.Tasks with a checkable right answer, such as math.
The three judge setups described in Zheng et al. (2023).

The same study is the reason the pattern caught on: strong judges agreed with human preferences more than 80% of the time, about as often as humans agree with each other. That is a statement about averages over many judgements. It does not make any single score trustworthy, and the same paper spends much of its length on where judges go wrong. Evaluation tools package the pattern — Langfuse runs LLM-as-a-judge evaluators on production observations and experiment datasets, and promptfoo’s llm-rubric assertion grades test outputs against a rubric — but the judge inside them has the same properties.

Which biases do LLM judges have?

BiasWhat happensWhat reduces it
PositionIn a pairwise comparison, the judge favours an answer for where it appears, not for what it says.Run the comparison in both orders; count a win only when it holds both times.
VerbosityThe judge favours longer answers, even when they are not clearer or more accurate.Rubric items that reward meeting the requirement, not covering more ground; length limits checked separately.
Self-preferenceThe judge scores its own outputs higher than others that humans rate as equal.Judge with a different model from the one that wrote the candidate; confirm against human labels.
Position and verbosity bias as defined in Zheng et al. (2023); self-preference as measured in Panickssery et al. (2024).

Self-preference is the subtle one. The MT-Bench authors saw some judges favour their own answers but could not establish the bias from their data. A later study set out to measure it directly:

One such bias is self-preference, where an LLM evaluator scores its own outputs higher than others' while human annotators consider them of equal quality.
— Panickssery et al., LLM Evaluators Recognize and Favor Their Own Generations (2024)

The same study found the bias grew with the model’s ability to recognise its own writing. That matters for any pipeline where one model both writes and grades — including, as described below, ours.

How do you make a judge’s scores consistent?

  • Use a coarse scale. Evidently AI’s judge guide notes that binary labels tend to be more reliable and consistent than fine-grained scores, for models and for people. If you need more than two levels, define every one.
  • One criterion per judge. When several aspects matter, Evidently recommends splitting them into separate evaluators rather than asking one call to weigh completeness, accuracy and tone at once.
  • Define every label. "Toxic" or "clear" means whatever the model guesses unless the rubric says what counts. Anchor each score with a description of what earns it.
  • Reason, then score. Evidently suggests asking the judge to explain its reasoning before it answers; G-Eval combines chain-of-thought evaluation steps with a form-filling score.
  • Constrain the answer. Ask for a fixed JSON shape and reject anything else, so a malformed reply fails loudly instead of being parsed into a number.
  • Validate against people. Hamel Husain’s method: have a domain expert label examples, treat the judge as a binary classifier, and measure its true positive and true negative rates — plain agreement hides a judge that misses rare failures.

Temperature is the setting people reach for first, and it helps less than expected. Tamba (2026) found that pinning temperature to 0 reduced pass/fail flips on borderline items but did not eliminate them across 690 API calls. And Claude 4.7 and later models reject a non-default temperature outright. Tamba recommends tracking grader disagreement across repeated runs; that, and a coarse scale, address variance that a sampling setting cannot — and one you may not be able to change.

Why is the text a judge grades an attack surface?

Because a judge reads the candidate as input, and a model cannot reliably tell text it is meant to evaluate from text that tells it what to do. If anyone can influence the candidate — a user’s prompt, a scraped page, another model’s answer — they can address the judge directly. This is not hypothetical: the JudgeDeceiver attack appends an optimized sequence to a candidate response so that the judge selects it regardless of what the other candidates say.

  • Mark the candidate as data. Wrap it in a named tag and tell the judge that the tagged text is the object of evaluation and that instructions inside it are not addressed to it. OWASP’s guidance on prompt injection says the same: separate and clearly denote untrusted content.
  • Bound the output. A strict schema and clamped values cannot stop a manipulated judge from choosing the highest allowed score, but they stop it from returning anything else.
  • Watch the top of the distribution. Inflated scores cluster at the maximum. Spot-check the highest-scored candidates, especially any that other users will see.

None of these makes injection impossible. Treat a judge’s score on text from an untrusted source as a signal to review, not as a verdict.

What does it cost to run a judge?

One more model call per score, with the rubric and the entire candidate as input. That sounds small until it is multiplied: a judge on every run of a pipeline adds a call to every run, and pairwise comparisons in both orders double it. The model choice drives the bill — prices differ by an order of magnitude between models, as Model Selection & Smart Routing covers.

VantagePrompt runs a judge on text, image, audio and video prompts, and on OpenRouter runs the judge call’s cost is part of the run’s credit charge. On free runs left on automatic model selection, the pipeline skips two of its three model calls, the intent classifier and the judge; those runs get a structural score instead. Which brings up the real question.

When is a heuristic score enough?

A heuristic is code that scores structure without reading meaning: does the output have the expected sections, is it long enough, does it contain lists. It is free, instant and perfectly repeatable. It is also blind — it cannot tell a precise prompt from a long, well-formatted vague one.

HeuristicLLM judge
Cost per scoreNoneOne model call
RepeatableExactlyMostly, with variance
Reads meaningNo — counts structureYes, through the rubric
Fails byRewarding length and formattingBias, drift, injection
Good forA floor: catching empty or malformed outputRanking candidates that pass the floor

The two do not share a scale, even when both print a number out of 100. In our pipeline the judge rates five dimensions from 1 to 5 and combines them, so its composite cannot fall below 20; the heuristic measures structure on its own scale and cannot see meaning at all. Never average one with the other, and never compare a heuristic score with a judge score as if they measured the same thing.

Use the heuristic as a gate, the judge to rank what passes it, and people — with the pass/fail method in How to Evaluate Prompt Quality — to decide what ships.

What does VantagePrompt’s judge actually do?

  • Rubric: five dimensions per output type (text, image, audio, video) — for text, clarity, specificity, structure, completeness and linguistic quality — each scored 1 (very poor) to 5 (excellent), weighted 0.25, 0.25, 0.20, 0.15 and 0.15, and scaled to a 100-point composite.
  • Strictness: the text judge is told not to inflate scores, that most prompts should average 3 to 4, and that only exceptional prompts earn 5s — which is why a competent prompt usually lands mid-range, as the quality score guide explains.
  • Constrained output: where the model supports it, the judge must answer in a strict JSON schema with every dimension required; each value is clamped to 1–5 before weighting.
  • Same model: by default, the judge runs on the model that generated the prompt. That is the setup self-preference research warns about, so we read our own score as a consistency check on the prompt’s structure, not as independent proof that it works.

Frequently asked questions

What is LLM-as-a-judge?
It is the practice of prompting a language model with a rubric to evaluate another model’s output, returning a score or a preference. The three common setups are grading a single answer, comparing two answers pairwise, and grading against a reference answer. It is fast and scalable, but it needs validating against human judgment.
What biases do LLM judges have?
Three are well documented. Position bias: in pairwise comparisons the judge favours an answer for where it appears. Verbosity bias: it favours longer answers even when they are not better. Self-preference: it scores its own outputs higher than others that humans rate as equal.
Should an LLM judge use temperature 0?
A low temperature reduces score flips but does not eliminate them, and Claude 4.7 and later models reject a non-default temperature. A coarse, fully defined scale, one criterion per judge, a fixed output schema and tracking disagreement across repeated runs address the variance a sampling setting cannot.
Can a judge model be manipulated by the text it scores?
Yes. The judge reads the candidate as input, so instructions inside the candidate can reach it; published attacks append optimized text that makes a judge pick a chosen answer. Mark the candidate as data, constrain the output to a schema, and review the highest scores on any text from an untrusted source.
Is a heuristic score as good as an LLM judge?
No, but it is often enough as a first gate. A heuristic counts structure — sections, length, lists — for free and repeatably, but cannot read meaning. Use it to reject empty or malformed output, a judge to rank what passes, and never compare or average scores from the two, because they are not on the same scale.

Sources

Put it into practice.

Run this technique in the optimizer.

Open the optimizer

Keep reading