All guides
QualityAdvanced·Intermediate7 min

The Prompt Quality Score, and How to Raise It

The quality score is a single 0-100 composite of clarity, specificity, structure, completeness, and linguistic quality, with clarity and specificity weighted heaviest. The judge is deliberately strict, so most competent prompts land mid-range; aim for "no ambiguity left" rather than for 100.

Every optimization returns a 0-100 quality score. Here's what that single number really measures, why a strict judge keeps most prompts mid-range, and how to pair it with opt-in deep analysis to spot subtle issues and iterate toward a sharper prompt.

By Andrei Bădulescu, Founder at VantagePrompt·Updated

Every optimization comes back with a quality score from 0 to 100. It is one number, not a stack of sub-scores. Read it right and it tells you whether a prompt is ready to ship or still leaking ambiguity. Read it wrong and you either over-polish a prompt that was already fine or trust one that quietly underperforms.

This guide covers what the number means, why a "good" prompt usually lands in the middle of the range, and how to use deep analysis to find the subtle issues a score alone won't name. Then you iterate.

What does the quality score actually measure?

The score is a composite. VantagePrompt judges your optimized output across several dimensions — clarity, specificity, structure, completeness, and linguistic quality — then weights and rolls them into one 0-100 figure. You see the total, not the breakdown. Clarity and specificity carry the most weight, so a prompt that is well-organized but vague about format, audience, or constraints will be capped by its weakest, heaviest dimensions.

The judge is deliberately strict. It is told not to inflate scores: most prompts are expected to average mid-rubric, and only genuinely exceptional prompts hit the top. That changes how you should read the band.

BandWhat it meansWhat to do
HighClear intent, concrete details, complete context.Ship it.
MidSolid, but missing specifics — format, audience, examples, or constraints. The usual home for a competent prompt.Add the missing specifics if the output is not landing.
LowVague, thin, or structurally broken.Rework before relying on it.
The judge is deliberately strict, so a mid-band score is the normal outcome, not a failure.

A mid-range score is not a failure. The judge reserves the top of the scale for prompts that leave nothing to interpretation. Chasing a perfect score is rarely worth it — aim for "no ambiguity left," not "100."

Why do two prompts for the same task score differently?

Our results reveal that strong LLM judges like GPT-4 can match both controlled and crowdsourced human preferences well, achieving over 80% agreement, the same level of agreement between humans.
— Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (2023)

Specificity is where most points are won or lost. Compare these two prompts for the same task:

Write a summary of this quarterly report for the team.
Low specificity — generic ask
Summarize this quarterly report for the engineering leads.
Output: 5 bullet points, each one sentence.
Focus on: shipped features, missed deadlines, and next-quarter risks.
Tone: factual, no marketing language. Skip financials.
High specificity — format, audience, constraints, length

The second names the audience, the output shape, the focus, and what to exclude. Those are exactly the things the specificity and completeness dimensions reward, which is why a tighter prompt lands higher even when both are grammatically clean.

What does deep analysis find that the score cannot?

The score tells you how good a prompt is. It does not tell you why it fell short. That is what deep analysis is for. It is an opt-in critic you trigger from the optimizer — a single pass that surfaces at most three subtle issues a simple linter would miss: semantic ambiguity, hidden assumptions, contradictory signals, scope drift, missing antecedents, or unclear success criteria.

Each issue comes back with a severity. A warning flags something that holds the prompt back; info is a nice-to-fix. Use the warnings as your edit list. A typical result looks like this:

{
  "issues": [
    {
      "severity": "warning",
      "label": "Ambiguous scope",
      "detail": "'The team' is undefined — engineering, sales, or all staff?"
    },
    {
      "severity": "info",
      "label": "No length bound",
      "detail": "Output length is unconstrained; add a sentence or word cap."
    }
  ]
}
Example deep-analysis output

Deep analysis costs 1 credit on a cache miss. Re-analyzing the same prompt within roughly a day is a cache hit and costs nothing — so iterate freely on one draft, but expect a charge when you analyze a fresh prompt.

What does the iteration loop look like?

The iteration loopOptimize, read the score, run deep analysis to name the specific weaknesses, edit the raw prompt resolving warnings first, then re-optimize. Stop when the warnings are gone and the score plateaus.OptimizeRead score0–100Deep analysis≤3 issuesEdit raw promptwarnings firstrepeat until the warnings are gone and the score plateausDo not grind for the last point.
Score tells you how good it is; deep analysis tells you why it fell short.

Put the score and deep analysis together into a tight cycle:

  1. Optimize and read the score. Low or mid? There's room to improve.
  2. Run deep analysis to name the specific weaknesses.
  3. Edit the raw prompt: resolve every warning first, then the info items worth fixing — usually adding format, audience, constraints, or examples.
  4. Re-optimize and compare the new score against the old one.
  5. Stop when the warnings are gone and the score plateaus. Don't grind for the last point.

For tasks that need step-by-step thinking, pair this with the reasoning use case and effort controls; for prompts that depend on current facts, grounding with web search closes the gaps deep analysis flags as hidden assumptions. The score keeps you honest; the loop is what moves it.

Frequently asked questions

What does the quality score actually measure?
It is one composite number, not a stack of sub-scores. VantagePrompt judges the optimized output across clarity, specificity, structure, completeness, and linguistic quality, then weights and rolls them into a single 0-100 figure. Clarity and specificity carry the most weight, so a well-organised but vague prompt is capped by its weakest heavy dimension.
Why is my score only mid-range when the prompt looks fine?
The judge is told not to inflate scores. Most prompts are expected to average mid-rubric and only genuinely exceptional prompts reach the top. A mid-range score is not a failure — chasing a perfect score is rarely worth it.
What is deep analysis, and how is it different from the score?
The score tells you how good a prompt is; it does not tell you why it fell short. Deep analysis is an opt-in critic that runs a single pass and surfaces at most three subtle issues a linter would miss: semantic ambiguity, hidden assumptions, contradictory signals, scope drift, missing antecedents, or unclear success criteria.
How much does deep analysis cost?
One credit on a cache miss. Re-analysing the same prompt within roughly a day is a cache hit and costs nothing, so you can iterate freely on one draft — but expect a charge when you analyse a fresh prompt.
How do I know when to stop iterating?
Stop when the warnings are gone and the score plateaus. Resolve every warning first, then the info items worth fixing, re-optimize, and compare. Do not grind for the last point.

Sources

Put it into practice.

Run this technique in the optimizer.

Open the optimizer

Keep reading