All guides
Audio/VideoMultimodal·Beginner7 min

Audio & Video Prompting: Suno, Runway, Sora & Veo

The Audio and Video use cases emit plain text in the target generator’s own grammar rather than XML, and they skip the prompt-engineering framework layer entirely, because RTF or CO-STAR scaffolding adds nothing to a song or a shot list. Pick a secondary tool selector — eight for audio, nine for video — and the same rough idea comes back shaped for Suno section markers or Runway camera direction.

VantagePrompt's Audio and Video use cases produce plain-text prompts tuned to a specific generator's grammar — not XML. Pick a tool selector and the optimizer shapes your rough idea into something Suno, Runway, Sora, or Veo can use directly.

By Andrei Bădulescu, Founder at VantagePrompt·Updated

Text prompts want structure. Audio and video prompts want grammar. A model writing a JSON schema cares about fields and types; Suno cares about [verse] and [chorus] markers; Runway cares about camera moves and shot duration. These are different languages, so VantagePrompt treats them differently.

When you pick the Audio or Video use case, the optimizer produces plain text tuned for the generator you name, not the XML structure used for text prompts. This guide covers both use cases, their tool selectors, and how to write inputs that come back as something you can paste straight into Suno, Runway, Sora, or Veo.

Why do audio and video prompts skip XML?

For the six textual use cases (auto, coding, reasoning, creative, marketing, structured), VantagePrompt emits an XML-structured prompt with tags like <role>, <task>, and <output_format>. That structure helps a downstream LLM parse intent reliably.

Generators like Suno, Runway, Sora, and Veo do not read XML. They expect their own native grammar. So the Audio and Video use cases bypass XML entirely and return plain text shaped for the target tool. The optimizer also skips the prompt-engineering framework layer for these use cases, because RTF or CO-STAR scaffolding adds nothing to a song or a shot list.

What audio tools can I target?

Select Audio, then choose a secondary tool. Each option steers the output toward that generator's conventions. The available tools:

Audio toolWhat the prompt is shaped around
universalTool-agnostic audio prompt, for when you have not settled on a generator.
elevenlabsVoice and speech direction — delivery, emotion, pacing.
sunoFull song prompts with section markers like [verse], [chorus], [bridge].
udioSong and music generation prompts.
barkExpressive speech and sound generation.
audioldmText-to-audio and sound design.
musicgenInstrumental and music generation.
stable_audioSound and music synthesis.
The Audio use case and its eight tool selectors.

The optimizer builds a tool-aware prompt for whichever you pick. For Suno it leans on section markers; for ElevenLabs it shapes voice and delivery. Wrap your raw idea in normal language and let the tool selector handle the grammar.

What does an optimized Suno prompt look like?

Use case: Audio. Tool: suno. Your raw input can be loose:

an upbeat indie-pop song about leaving a small town
at the end of summer, bittersweet but hopeful, female vocals

The optimizer returns plain text built around Suno's structure — a style/genre line plus section markers you can drop straight in:

[Style: indie pop, bright guitars, mid-tempo, female lead vocal]

[Verse]
Gravel road and a half-packed car...

[Chorus]
This town gets smaller in the mirror...

[Bridge]
...

[Outro]
...

Put the emotional intent and any hard constraints (vocal gender, tempo, instruments, language) in your raw input. The optimizer is good at structure; it cannot guess the mood you never stated.

What video tools can I target?

Select Video, then a tool. Video generators differ sharply in how they take direction — camera language, shot length, motion cues — so the secondary selector matters even more here. Available tools:

Video toolWhat the prompt is shaped around
universalGenerator-agnostic video prompt.
runwayGen-style prompts with camera and motion direction.
pikaShort-form motion and scene prompts.
soraNarrative scene description prompts.
klingCinematic motion prompts.
lumaDream Machine-style prompts.
stable_videoImage-to-video and motion prompts.
minimaxScene and motion prompts.
veoGoogle Veo scene and cinematography prompts.
The Video use case and its nine tool selectors — direction language differs sharply between them.
Use case: Video  ·  Tool: runway

Raw input:
a lone lighthouse on a cliff during a storm, waves crashing,
slow push-in, moody and cinematic, golden hour breaking through clouds

Because the tool is set to Runway, the returned plain text emphasizes camera movement, subject, lighting, and atmosphere in the order Runway responds to best — rather than an abstract paragraph. Switch the tool to Sora and the same idea comes back framed as a narrative scene description instead.

How do I get sharper audio and video prompts?

  • Always set the secondary tool when you know your target — universal is a safe default, but a named tool produces sharper, paste-ready output.
  • State constraints explicitly: aspect ratio intent, duration, vocal gender, instruments, tempo, mood. The optimizer structures what you give it; it does not invent unstated preferences.
  • Audio and video output is plain text, so copy it directly into the generator — there are no XML tags to strip.
  • Switching tools re-shapes the same idea for a different grammar, so try the same raw input across two tools when you are deciding which generator to use.
  • Like every run, audio and video optimizations are metered in credits and you get a quality score on the result.

Image generation has its own use case and tool selectors (Midjourney, DALL·E, Flux, and more) — see the Image Prompting guide. If you are unsure whether your task is audio, video, or something textual, the Choosing the Right Use Case guide walks through all of them.

Frequently asked questions

Why do audio and video prompts skip XML?
Generators like Suno, Runway, Sora, and Veo do not read XML — they expect their own native grammar. The Audio and Video use cases therefore return plain text shaped for the target tool, and also skip the prompt-engineering framework layer, because RTF or CO-STAR scaffolding adds nothing to a song or a shot list.
Do I have to pick a tool selector?
No, but you should when you know your target. Universal is a safe default; a named tool produces sharper, paste-ready output. Switching tools re-shapes the same idea for a different grammar, which makes it easy to compare two generators from one raw input.
What should I put in the raw input for a song or a shot?
The emotional intent and every hard constraint: vocal gender, tempo, instruments, language, aspect-ratio intent, duration, mood. The optimizer is good at structure but cannot guess a preference you never stated.
Do audio and video runs cost credits and get a quality score?
Yes — like every run, they are metered in credits and the result carries a quality score.
Where does image generation fit?
Image has its own use case and its own tool selectors — Midjourney, DALL·E, Flux and more — covered in the Image Prompting guide. It follows the same plain-text-per-tool pattern.

Sources

Put it into practice.

Run this technique in the optimizer.

Open the optimizer

Keep reading