Every AI image tool — ChatGPT, Midjourney, Gemini, Stable Diffusion — speaks the same underlying language: visual description. The tools differ in dialect (parameters here, negative fields there), but the core skill is one skill. Learn it once, use it everywhere. That’s this guide.
Below: the universal formula and its seven slots, the word banks that fill them, how the same prompt adapts across tool dialects, recipes organized by what you’re actually making, and fixes for the problems every tool shares. Tool-specific depth lives in our Midjourney guide; this is the layer above it. Guide #24 of the Prompt Engineering roadmap.
How Image Models “Think” (30 Seconds)
Image models learned by pairing millions of images with their descriptions. When you prompt, the model maps your words to the visual concepts those words appeared with — your vocabulary literally selects from what it learned. “Beautiful” appeared with everything, so it selects nothing. “Golden hour rim lighting” appeared with a very specific look, so it delivers exactly that. This is why image prompting is a vocabulary game: concrete visual words are the only controls you have.
Quick Start — Feel the Difference in 2 Minutes
Open any image tool and run this pair — the fastest lesson in this entire guide:
→ same subject, different universe
Nothing about round 2 required talent — just filled slots. That’s the promise of the formula below: image quality as a checklist, not a gift.
The Universal Formula — 7 Slots
What each slot buys you
- 1. Subject + one detail: “a fox” is a lottery ticket; “a red fox with snow on its muzzle” is a commission. One vivid detail anchors everything.
- 2. Action/pose: verbs create life — sitting, leaping, mending, glancing back. Static subjects read as stock photos.
- 3. Environment: where + when. “Forest clearing at dusk” does more work than three adjectives.
- 4. Style/medium: the single biggest lever — photograph vs watercolor vs flat vector vs 3D render are different universes.
- 5. Lighting: the 80/20 of image quality — see the mini-bank below.
- 6. Mood: one emotional word tunes color and atmosphere: serene, ominous, whimsical, nostalgic.
- 7. Composition: close-up / wide shot / centered / rule of thirds / from above / empty space left for text.
- Exclusions: what must NOT appear — full craft in Negative Prompting.
Word Banks — Fill the Slots Fast
- Styles: photorealistic · documentary photography · cinematic still · flat vector · watercolor · oil painting · ink sketch · isometric · low-poly 3D · pixel art · paper cutout · claymation
- Lighting: golden hour · blue hour · soft window light · dramatic side light · volumetric rays · neon glow · overcast diffused · candlelit · backlit silhouette
- Composition: extreme close-up · medium shot · wide establishing shot · bird’s-eye view · low angle · centered symmetrical · rule of thirds · shallow depth of field · empty upper third for title
- Mood: serene · triumphant · melancholic · cozy · ominous · playful · majestic · nostalgic
One Prompt, Every Tool — The Dialect Table
The description is universal; the exclusions and settings change costume per tool:
Practical upshot: write your prompt once in formula form, then translate — sentences for chat tools, phrases+params for MJ, keywords+negative field for SD. Same image intent, three costumes.
One Subject, Five Styles — Watch Slot 4 Work
To burn in how much the style slot matters, here’s one identical scene through five media — steal whichever matches your project:
- + “dramatic cinematic photography, long exposure” → movie-poster realism
- + “flat vector illustration, limited palette of navy and cream” → modern website hero
- + “children’s book watercolor, soft edges, gentle colors” → storybook warmth
- + “isometric low-poly 3D render, pastel colors” → game-art charm
- + “vintage travel poster, bold shapes, grainy texture, muted reds” → retro print
Five publishable directions, one sentence changed. When a client or a blog post needs “options,” this is how you generate them in minutes.
Recipes by What You’re Making
Iterating on Images — The Rules
- One variable per generation. Change lighting OR style OR angle — never all three. Otherwise you can’t learn what caused what (refinement rules, visual edition).
- Reroll with intent, not hope: if two consecutive generations disappoint the same way, the prompt is the problem — name the failure (too busy? wrong mood?) and fix that slot instead of rerolling a broken recipe
- Anchor before exploring: get subject + composition right first in a plain style, THEN try style variations — style experiments on a broken composition waste every generation.
- Vary the winner, not the prompt: when a result is close, use the tool’s variation feature on THAT image instead of rerolling from scratch.
- Save winning prompts verbatim — your visual recipe book is the asset; a hit prompt with [bracketed] subject becomes a reusable template.
Universal Problems & Fixes
Mistakes That Waste Generations
- Empty adjectives as strategy: beautiful, stunning, amazing, 8k, masterpiece — training saw them everywhere, so they select nothing; slots and word banks are the actual controls
- Instruction-speak: “please create an image that shows…” — the polite wrapper is dead weight in every tool; describe the finished picture
- Changing everything each round: new style + new light + new angle at once teaches you nothing about what worked — one variable, always
- Generating first, choosing format later: composition dies in the crop — pick orientation for the destination before the first generation
- Fighting text generation: spending ten rounds trying to fix garbled lettering that an editor would add perfectly in thirty seconds
- No recipe book: every winning prompt you don’t save is a future hour you’ll spend rediscovering it — the library habit applies to images doubly
Ethics & Rights — The 60-Second Version
- Real people: don’t generate identifiable real individuals without consent — deepfake territory, legally and ethically
- Living artists & trademarks: describe style qualities (“bold ink lines, muted palette”) instead of prompting names/characters for commercial work
- Commercial rights vary by tool and plan — check current terms; AI-image copyright law is still evolving and differs by country
- Disclose where it matters: product photos that misrepresent, news-like images — transparency isn’t just ethics, it’s increasingly law
Which Tool for Which Job?
- Quick + conversational: ChatGPT/Gemini images — iterate by chatting, no syntax to learn (ChatGPT guide · Gemini guide)
- Aesthetic ceiling + control dials: Midjourney — the artist’s pick
- Free, local, infinitely tweakable: Stable Diffusion family — the tinkerer’s pick, steepest curve
- Honest answer: the formula matters more than the tool — a skilled prompt in a free tool beats a lazy prompt in the best one
Frequently Asked Questions
Subject with one defining detail + action + environment + style/medium + lighting + mood + composition + exclusions. Fill 4–5 of those slots with concrete visual words and you’re ahead of 90% of prompts in any tool.
The descriptive core transfers completely; only the costume changes — sentences for chat tools, phrases plus –parameters for Midjourney, keywords plus a negative field for Stable Diffusion.
Empty adjectives. ‘Beautiful, stunning, high quality’ select nothing because they described everything in training. Concrete medium, lighting, and mood words are the actual controls.
Keep it to 1–3 words at most — or generate the image clean and add typography in Canva/Photoshop. For anything professional, added text beats generated text every time.
Two or three that agree with each other. Stacking ten descriptors makes elements fight and produces muddy results — coherence beats quantity.
Fix a style phrase and palette and reuse them verbatim; use seed/reference features where the tool offers them; and change only the subject slot between generations.
Usually yes, subject to each tool’s plan terms — verify commercial rights, avoid trademarked content and real people, and remember copyright law for AI images is still settling and varies by country.
Lighting. One specific lighting phrase (golden hour, soft window light, dramatic side light) transforms an image more than any other single addition.
For best results, yes — most models were trained predominantly on English captions, so English visual vocabulary maps most reliably. Chat-based tools can translate your idea first: describe in your language, ask for an English image prompt, then run it.
This is the universal layer — the formula, word banks, and fixes that work in every tool. The Midjourney guide adds MJ’s specific dialect: parameters, workflow buttons, and advanced features like weights and permutations.
One formula. Every tool. 🖼️
Next: AI Video Prompting Guide — guide #25.




