All guides
ReasoningAdvanced·Intermediate7 min

Chain-of-Thought: Make Models Show Their Work

Chain-of-thought prompting asks a model to decompose a problem and reason through each step before it commits to an answer, which replaces a confident guess with a derivation you can audit. Use it for multi-step math, logic, debugging, and planning; skip it for lookups, rewrites, and formatting, where it only adds latency and cost.

Hard, multi-step questions are where models cut corners. Chain-of-thought makes them slow down and show their work. Here's when to reach for it and how to dial it in with VantagePrompt's Reasoning use case and reasoning-effort control.

By Andrei Bădulescu, Founder at VantagePrompt·Updated

Hard problems break models. Ask for a one-line answer to a multi-step question and you get a confident guess that skips the middle. Chain-of-thought fixes this: you tell the model to reason in steps before it commits to an answer. The work becomes visible, the logic becomes checkable, and the failure modes move from silent to obvious.

VantagePrompt has a use case built for exactly this. Pick Reasoning, dial in how much thinking budget you want, and the optimizer produces a prompt engineered to make the model show its work. This guide covers when chain-of-thought actually helps, how to trigger it, and how the reasoning-effort control changes the output.

What is chain-of-thought prompting?

Chain-of-thought (CoT) is a prompting technique: instead of asking for a final answer, you ask the model to decompose the problem, reason through each part, and only then conclude. The intermediate steps are not just for show. They constrain the model — each step has to follow from the last — so the final answer is anchored to a visible derivation instead of a snap judgment.

It is most useful when the answer depends on a chain of dependent sub-conclusions: math, multi-constraint logic, debugging, planning, trade-off analysis, anything where a single skipped step poisons the result.

When does chain-of-thought help, and when does it hurt?

We explore how generating a chain of thought — a series of intermediate reasoning steps — significantly improves the ability of large language models to perform complex reasoning.
— Wei et al., Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (2022)
  • Use it for multi-step math, logic puzzles, root-cause analysis, architecture trade-offs, and any task where the answer needs justification.
  • Use it when you need to audit the model's logic — visible steps let you catch a wrong assumption before you trust the conclusion.
  • Skip it for lookups, simple rewrites, formatting, or short factual answers. Forcing reasoning onto a trivial task just adds latency and token cost for no quality gain.
  • Skip it when you need terse machine-parseable output and nothing else — reasoning prose and a clean JSON payload pull in opposite directions.

How do I trigger chain-of-thought in VantagePrompt?

On the optimizer, set the use case to Reasoning. This optimization path is tuned for chain-of-thought, step-by-step decomposition, and self-verification. Picking it also skips the intent-classifier step that the default "auto" use case runs — you are telling the app the intent directly, which saves a round-trip. (For how the other use cases differ, see Choosing the Right Use Case.)

The optimizer returns an XML-structured prompt. For reasoning work the load-bearing tags are <instructions>, which lays out the step-by-step procedure and ordering, and <verification>, an auto-generated self-check the downstream model runs before finalizing. For reasoning prompts that check includes whether claims are supported, whether counterarguments were considered, and whether the conclusion actually follows from the steps.

You don't write the XML or the verification checklist yourself. You paste a rough question; VantagePrompt expands it into the structured prompt. You then run that prompt against your target model to get the reasoned answer.

What does the reasoning-effort control change?

Separately from the use case, the optimizer exposes a reasoning-effort setting with five levels: minimal, low, medium, high, and xhigh. This maps to the target model's reasoning budget — how much internal thinking it is allowed to spend before answering. Higher effort means more deliberation (and more tokens, which cost more credits); lower effort means faster, cheaper, shallower.

EffortReach for it whenCost & latency
minimal / lowQuick decompositions and simple multi-step tasks where a light pass is enough.Cheapest, fastest.
mediumThe practical default for most reasoning work — real step-by-step without a large budget.Moderate.
high / xhighGnarly proofs, deep debugging, multi-constraint planning where you want the model to grind.Highest latency and credit cost.
Reasoning-effort levels and what each is worth paying for.

Match the dial to the difficulty. xhigh on a two-step question wastes credits; minimal on a hard proof gives you the same shallow guess you were trying to avoid.

What does an optimized reasoning prompt look like?

A rough input on the left, the kind of structured, reasoning-first prompt the optimizer produces on the right. Note the explicit step ordering and the self-check tag.

Three services share one database. Latency spiked at 9am but only on writes. CPU is flat. What's the likely cause and how do I confirm it?
Rough input — pasted into the optimizer with use case = Reasoning, effort = high
<role>Senior systems engineer diagnosing a production incident.</role>
<task>Identify the most likely cause of a write-only latency spike and give a confirmation plan.</task>
<instructions>
  1. List candidate causes that fit write-only latency with flat CPU.
  2. For each candidate, state what evidence would confirm or rule it out.
  3. Rank candidates by likelihood given the symptoms.
  4. Give the single highest-probability cause and the exact check to confirm it.
</instructions>
<verification>
  - Each ranked claim is supported by a stated symptom.
  - At least one counter-hypothesis was considered and addressed.
  - The final conclusion follows from the ranking, not asserted independently.
</verification>
Optimized prompt (abridged) — paste this into your target model

The model can no longer jump to "it's a slow query." It has to enumerate, weigh evidence, and justify the pick — and the verification tag makes it check that the conclusion actually follows before it stops.

How do I improve a reasoning prompt that falls short?

Once you have a reasoning prompt, refine it. If the chain skips a step, tighten the question and regenerate. If the reasoning is sound but the answer feels thin, raise the effort level. To catch subtle problems — a hidden assumption, an ambiguous constraint — and to read the quality score the optimizer assigns, see Reading the Quality Score & Iterating.

Frequently asked questions

When should I not use chain-of-thought?
Skip it for lookups, simple rewrites, formatting, and short factual answers — forcing reasoning onto a trivial task adds latency and token cost with no quality gain. Skip it too when you need terse machine-parseable output and nothing else, because reasoning prose and a clean JSON payload pull in opposite directions.
What does the reasoning-effort setting actually change?
It maps to the target model's reasoning budget — how much internal thinking it is allowed to spend before answering. The five levels are minimal, low, medium, high, and xhigh. Higher effort means more deliberation and more tokens, which cost more credits; lower effort is faster, cheaper, and shallower. Medium is the practical default.
Do I have to write the XML and the verification checklist myself?
No. You paste a rough question and VantagePrompt expands it into the structured prompt, including the <instructions> step ordering and the auto-generated <verification> self-check. You then run that prompt against your target model to get the reasoned answer.
Does picking the Reasoning use case change how the pipeline runs?
Yes. Selecting Reasoning skips the intent-classifier step that the default auto use case runs, because you are telling the app the intent directly. That saves a round-trip on every optimization.
My reasoning chain skips a step — how do I fix it?
Tighten the question and regenerate. If the reasoning is sound but the answer feels thin, raise the effort level instead. The quality score the optimizer assigns tells you whether the change actually helped.

Sources

reasoning prompts to try

Browse all reasoning prompts

Published by the community and free to copy — worked examples of what this guide describes.

Put it into practice.

Run this technique in the optimizer.

Open the optimizer

Keep reading