All guides
EvaluationProduction·Intermediate8 min

How to Evaluate Prompt Quality by Its Outputs

You evaluate a prompt by its outputs, not by reading it: run it several times on real inputs at the settings you will use, check every output against pass/fail criteria written beforehand, and judge the prompt by how often it passes — not by its best answer.

A prompt can read beautifully and still fail one time in five. Reading it tells you how it looks; only its outputs tell you how it behaves. This is the manual method — no tooling, just a model, a handful of inputs and a checklist — for deciding whether a prompt is good enough, whether a rewrite is better, and when to stop editing.

By Andrei Bădulescu, Founder at VantagePrompt·Updated

Why can’t you judge a prompt by reading it?

The most common mistake in prompt evaluation is grading the prompt as a piece of writing. A prompt with clean sections, a role and a format block looks finished, and it is tempting to call it good on sight. But the reader that matters is the model, and it does not grade prose. A tidy prompt can leave the one decision that matters unmade; a clumsy one can pin it down exactly.

So treat every prompt as a hypothesis — "this text makes the model produce what I need" — and test it the way you would test any hypothesis: against outcomes, on more than one trial. Everything below is how to run that test by hand.

Why run the same prompt more than once?

Because the same prompt does not produce the same output twice. Sampling is random by design, and even the settings meant to remove that randomness do not fully succeed: OpenAI states that determinism is not guaranteed even with a fixed seed, and Thinking Machines Lab notes that LLM APIs are not deterministic in practice even at temperature 0. OpenAI’s own evaluation guide names this as the reason traditional software testing is not enough for model output.

One run tells you what the prompt can do. Several runs tell you what it usually does — and production only ever sees "usually". Run each input five times at exactly the settings you will use in production, and judge the spread, not the best output. If four answers are excellent and one ignores the format, the prompt has a one-in-five format failure, and that is the fact to record.

Five is not a statistical threshold; it is a practical number at which a one-in-five failure shows up about two times in three. A prompt that fails one run in five, tested once, looks perfect four times out of five:

Runs per inputChance of seeing a 1-in-5 failure at least once
120%
349%
567%
1089%
Probability of at least one failure in n runs when each run fails independently with probability 0.2: 1 − 0.8ⁿ.

Do not "fix" variance by lowering the temperature for the test and raising it again in production. Test at the settings you ship with. Claude 4.7 and later models reject a non-default temperature outright, which makes the point for you.

Use three to five real inputs, not one. Include the awkward one — the long document, the ambiguous request, the input that broke the last version. Three inputs × five runs is fifteen outputs per prompt version: enough to see a pattern, few enough to read every one.

What should you score in each output?

Write the checks down before you run anything — Anthropic’s evaluation docs start from exactly this: define success criteria that are specific and measurable, then design the test around them. Criteria written after you have seen the outputs bend towards whatever the outputs happened to do. Five dimensions cover most text tasks:

DimensionThe questionA pass/fail check
Instruction adherenceDid it do what was asked, and not something adjacent?Answers the question asked; respects every stated exclusion.
Factual groundingIs every claim supported by the input or by checkable fact?No figure, name or quote that is absent from the source.
Format complianceIs the shape exactly the one requested?Parses as JSON / has the five columns / stays under 150 words.
CompletenessIs anything the task required missing?Covers all three risks named in the brief.
Tone fitWould the intended reader accept it as written?No marketing language in an internal engineering summary.
Turn each dimension into one or more concrete checks for your task before the first run.

Five is a working limit, not a law. Every extra dimension is another judgement per output, and fifteen outputs × eight vague dimensions is where people start skimming. If a dimension does not matter for the task — tone, for a JSON extractor — drop it rather than scoring it out of habit.

Why use pass/fail checks instead of a 1–10 score?

Because you will not give the same output the same 7 twice, and neither will a colleague. A graded score asks for a fine distinction nobody has defined — what separates a 6 from a 7? A binary check asks a question with an answer: does it parse, is the exclusion respected, is the number in the source. Evidently AI’s guide to LLM judges makes the same recommendation, for human graders as well as model graders:

Binary evaluations, like "Polite" vs. "Impolite," tend to be more reliable and consistent for both LLMs and human evaluators.
— Evidently AI, LLM-as-a-judge: a complete guide

Hamel Husain argues the practical side: a 3 or a 4 tells nobody what to fix, while a failed check names the problem. Graded scales are not wrong — Anthropic’s docs list Likert and ordinal scales alongside binary ones — but for a person scoring fifteen outputs by hand, pass/fail is the version you will apply the same way on the fifteenth output as on the first. If a check is genuinely partial, split it into two binary checks rather than reaching for a scale.

The score for a prompt version is then a pass rate per check: format 15/15, grounding 13/15, completeness 15/15. That tells you exactly where the prompt leaks.

How do you compare two versions of a prompt without fooling yourself?

The paired test in the prompt optimization guide is the outline: same inputs, same model, same settings, criteria written first. It asks for a second run when outputs vary; that tells you a prompt varies, while five runs tell you how often it fails. The part that goes wrong in practice is the comparison itself, because you know which version you wrote last and you want it to win.

  1. Run both versions on the same three to five inputs, five runs each, at identical settings.
  2. Hide which version produced which output. Paste them into one document under neutral labels, or have someone else shuffle them.
  3. Score every output against the same pass/fail checks, one check at a time across all outputs, rather than one output at a time across all checks.
  4. When you compare two outputs side by side, read them in both orders and count a win only when one is preferred both times.
  5. Compare pass rates per check. Treat a gap of one or two passes out of fifteen as a tie — at this sample size gaps that small are noise, and even a larger gap deserves a re-run before you trust it.

Step 4 exists because order changes judgement. It was documented for LLM judges in the MT-Bench study, which also notes the same bias in human decision-making, and the same fix is easy to apply when you are the one reading two outputs:

A conservative approach is to call a judge twice by swapping the order of two answers and only declare a win when an answer is preferred in both orders. If the results are inconsistent after swapping, we can call it a tie.
— Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (2023)

A longer output is not a better output. A rewritten prompt often produces longer answers, and length reads as effort. Score against the checks, never against the word count.

When should you stop optimizing a prompt?

Neither OpenAI’s nor Anthropic’s evaluation docs give a stopping rule, so this one follows from the method rather than from a source. If you wrote the criteria first, the finish line is already defined:

  • Every check passes on every run on every input. The prompt meets the bar you set. Ship it, and keep the inputs and checks for the next model change.
  • Two edits in a row move no pass rate. You are polishing wording the model does not respond to. Stop editing the prompt and look at what is left.
  • The remaining failures are not the prompt’s to fix. A model that invents a figure because the source is missing needs the source, not another instruction — that is a grounding problem. A model that cannot hold a twelve-step procedure may need a different model.

Keep the test set. A prompt that passes today can drift when the model underneath is updated, and re-running fifteen outputs against checks you already wrote takes minutes.

Where does a built-in quality score fit?

VantagePrompt scores every optimized prompt from 0 to 100. For text prompts that score normally comes from a judge model rating the optimized prompt itself — clarity, specificity, structure, completeness and linguistic quality — though free runs left on automatic model selection, or a run where the judge fails, get a faster structural check instead. Either way it is a read on the prompt’s text, before any output exists. Use it to rank drafts and catch an obviously underspecified one; then run the output test on the drafts worth keeping. The quality score guide covers how to read the number and how opt-in deep analysis names up to three specific gaps.

If you have several drafts to compare, a batch optimizes 2 to 20 prompts in one submission. And if you find yourself running this test by hand more than a few times a week, automate the mechanics: a harness such as promptfoo can run each test case several times with its --repeat option, and a model can apply the same pass/fail checks — at which point the biases above become the judge’s to manage rather than yours. LLM-as-a-Judge covers how to run one and what it gets wrong.

What does a one-page evaluation checklist look like?

  • Pass/fail checks written down before the first run.
  • Three to five real inputs, including the one that usually breaks things.
  • Five runs per input, at production settings.
  • Every output read and scored, one check at a time.
  • A pass rate per check — not an average of everything.
  • Versions compared blind, in both orders, with small gaps called ties.
  • A stop when the checks pass, or when edits stop moving them.

Frequently asked questions

How do I evaluate the quality of a prompt?
Judge it by its outputs, not by reading it. Write pass/fail checks for what a correct answer must contain, run the prompt several times on three to five real inputs at the settings you will use, score every output against the checks, and record a pass rate per check. The prompt is as good as how often it passes, not as good as its best answer.
How many times should I run a prompt to test it?
Five runs per input is a practical minimum. Model output varies between runs even at temperature 0, and a prompt that fails one run in five shows a failure in only about half of three-run tests but in about two thirds of five-run tests. Use three to five inputs, so one version gets roughly fifteen outputs.
Is a 1–10 score or a pass/fail check better for evaluating prompts?
For manual evaluation, pass/fail. People rarely give the same output the same number twice, and a 6 or a 7 does not say what to fix, while a failed check names the problem. Graded scales are valid, but binary checks are easier to apply consistently across fifteen outputs.
How do I A/B test two versions of a prompt?
Run both on the same inputs, the same model and the same settings, several times each. Hide which version produced which output, score all outputs against the same pass/fail checks, read side-by-side pairs in both orders, and compare pass rates per check. Treat a difference of one or two passes out of fifteen as a tie.
When should I stop optimizing a prompt?
When every pre-written check passes on every run and input, or when two consecutive edits fail to change any pass rate. If the remaining failures come from missing information or a model limit, more prompt edits will not fix them: supply the source, or try a different model.

Sources

Put it into practice.

Run this technique in the optimizer.

Open the optimizer

Keep reading