All guides
ImageMultimodal·Beginner6 min

Image Prompting: Midjourney, DALL·E, Flux & More

The Image use case emits plain text rather than the XML that textual use cases produce, because image generators each have their own grammar — Midjourney flags, DALL·E 3 scene prose, dense Flux description. Pick one of eight tool selectors and the output is shaped for that exact generator; picking Image also skips the intent-classifier step.

Image generators speak their own dialects — flags for Midjourney, fluent scenes for DALL-E 3, dense description for Flux. VantagePrompt's Image use case skips XML entirely and emits plain text shaped for the exact generator you select.

By Andrei Bădulescu, Founder at VantagePrompt·Updated

Image generators do not read XML. Midjourney wants comma-separated phrases and flags like --ar 16:9. DALL-E 3 wants a fluent natural-language scene. Flux likes dense, layered description. A prompt that works for one is mediocre on another. The Image use case in VantagePrompt knows this: it skips the XML structure entirely and emits plain text shaped for the exact generator you are targeting.

This guide covers the Image use case and its tool selectors: what each does, why image output is plain text instead of the tagged XML you get for coding or structured prompts, and how to turn a one-line idea into a generator-ready prompt.

Why is image output plain text and not XML?

For textual use cases, VantagePrompt produces a structured XML prompt with tags like <role>, <task>, and <output_format> — a format downstream LLMs parse reliably. Image generators do not work that way. They each have their own grammar: aspect-ratio flags, style tokens, weighting syntax. A structured tag scheme would just get in the way.

So when you pick the Image use case, the optimizer routes to a dedicated generator that returns trimmed plain text — no XML wrapper, no structural tags. The OutputPanel renders it as plain text instead of syntax-highlighted XML. Paste it straight into your image tool.

The Image use case also skips the intent-classifier step that textual prompts run. You are telling the app explicitly what you want, so it does not need to guess — that saves a step and gets you to the result faster.

What does each image tool selector change?

After choosing the Image use case, a secondary selector appears. Pick the generator you are actually going to use — the output grammar is tuned for that target.

SelectorPrompt grammar it emits
universalModel-agnostic phrasing — for when you are not committed to one generator, or want a prompt you can adapt by hand.
midjourneyComma-separated descriptors plus flags like --ar (aspect ratio), --style raw, and --v (version).
dalle3Fluent natural-language scene description. DALL·E 3 rewrites terse prompts, so explicit beats keyword-soup.
stable_diffusionDense descriptive phrasing, the style most Stable Diffusion workflows expect.
fluxLayered, detailed prompting that plays to Flux's strengths.
leonardoPhrasing tuned for Leonardo's pipeline.
imagen2Phrasing for Google's Imagen 2.
openai_imageOpenAI's image endpoint.
The selector only applies while the Image use case is active.

The selector is only meaningful when the Image use case is active. If you switch use cases, it is ignored. Default to universal when in doubt, then re-run with a specific tool once you know your generator.

What does an optimized image prompt look like?

Start with a rough idea — the same input you would type into the optimizer's textarea:

a cozy coffee shop on a rainy evening, warm light

With Image + midjourney selected, the optimizer expands that into a generator-ready plain-text prompt with concrete composition, lighting, and the flags Midjourney expects — something along these lines:

cozy corner coffee shop interior at dusk, rain streaking the front
windows, warm amber pendant lights and a glowing espresso machine,
steam rising from a fresh cup, reflections on a wet wooden counter,
shallow depth of field, moody cinematic atmosphere --ar 16:9 --v 6

Pick dalle3 instead and the same idea comes back as a flowing natural-language paragraph with no flags — because that is what DALL-E 3 responds to. The idea is identical; the grammar changes with the target.

How do I get better results from the Image use case?

  • Give the optimizer a subject and a mood, not just a noun. "A dragon" has little to work with; "a weathered red dragon coiled on a cliff at sunset, dramatic backlight" gives it real material to expand.
  • Match the selector to your actual generator before you copy. Midjourney flags will read as literal text in DALL-E 3 and quietly degrade the result.
  • Use universal first if you are comparison-shopping generators, then regenerate per tool once you commit.
  • Every run still gets a quality score, so you can judge the prompt before spending a generation credit downstream. Each optimization run is metered in credits like any other use case.

Targeting audio or video generators instead? Suno, Runway, Sora, and Veo follow the same plain-text-per-tool pattern — see the Audio & Video Prompting guide. Not sure the Image use case is even the right pick? Start with Choosing the Right Use Case.

Frequently asked questions

Why is image output plain text instead of XML?
Image generators do not read XML. Each has its own grammar — aspect-ratio flags, style tokens, weighting syntax — and a structured tag scheme would get in the way. The Image use case routes to a dedicated generator that returns trimmed plain text with no XML wrapper, which you paste straight into your image tool.
Which image tool selector should I choose?
The one you will actually generate with, because the grammar differs: Midjourney flags read as literal text in DALL·E 3 and quietly degrade the result. Default to universal when comparison-shopping generators, then regenerate per tool once you commit.
Does the Image use case skip any pipeline steps?
Yes — it skips the intent-classifier step that textual prompts run. You are stating explicitly what you want, so nothing has to be guessed, which saves a round-trip.
Do image runs still get a quality score and cost credits?
Yes. Every run gets a quality score, so you can judge the prompt before spending a generation credit downstream, and each optimization is metered in credits like any other use case.
What makes a good input for the Image use case?
Give a subject and a mood, not just a noun. "A dragon" has little to work with; "a weathered red dragon coiled on a cliff at sunset, dramatic backlight" gives the optimizer real material to expand.

Sources

image prompts to try

Browse all image prompts

Published by the community and free to copy — worked examples of what this guide describes.

Put it into practice.

Run this technique in the optimizer.

Open the optimizer

Keep reading