Prompt Engineering

AI Image Prompting Guide 2026 Works in Every Tool

AI Image Prompting Guide 2026 - Techprofree

Every AI image tool — ChatGPT, Midjourney, Gemini, Stable Diffusion — speaks the same underlying language: visual description. The tools differ in dialect (parameters here, negative fields there), but the core skill is one skill. Learn it once, use it everywhere. That’s this guide.

Below: the universal formula and its seven slots, the word banks that fill them, how the same prompt adapts across tool dialects, recipes organized by what you’re actually making, and fixes for the problems every tool shares. Tool-specific depth lives in our Midjourney guide; this is the layer above it. Guide #24 of the Prompt Engineering roadmap.

How Image Models “Think” (30 Seconds)

Image models learned by pairing millions of images with their descriptions. When you prompt, the model maps your words to the visual concepts those words appeared with — your vocabulary literally selects from what it learned. “Beautiful” appeared with everything, so it selects nothing. “Golden hour rim lighting” appeared with a very specific look, so it delivers exactly that. This is why image prompting is a vocabulary game: concrete visual words are the only controls you have.

Quick Start — Feel the Difference in 2 Minutes

Open any image tool and run this pair — the fastest lesson in this entire guide:

❌ ROUND 1: “a beautiful cat” — accept whatever generic result appears
✅ ROUND 2: “a fluffy ginger cat curled on a windowsill, rainy evening outside, soft warm lamp light, cozy nostalgic mood, watercolor illustration, close-up — no text”
→ same subject, different universe

Nothing about round 2 required talent — just filled slots. That’s the promise of the formula below: image quality as a checklist, not a gift.

The Universal Formula — 7 Slots

THE FORMULA (WORKS EVERYWHERE)[1 SUBJECT + defining detail] + [2 ACTION/POSE] + [3 ENVIRONMENT] + [4 STYLE/MEDIUM] + [5 LIGHTING] + [6 MOOD] + [7 COMPOSITION] + [exclusions]
FILLED“a young chess player concentrating [1+2], in a sunlit library [3], watercolor illustration [4], soft morning light through tall windows [5], quiet focused mood [6], medium shot slightly from above [7] — no text, no clutter [exclusions]”

What each slot buys you

  • 1. Subject + one detail: “a fox” is a lottery ticket; “a red fox with snow on its muzzle” is a commission. One vivid detail anchors everything.
  • 2. Action/pose: verbs create life — sitting, leaping, mending, glancing back. Static subjects read as stock photos.
  • 3. Environment: where + when. “Forest clearing at dusk” does more work than three adjectives.
  • 4. Style/medium: the single biggest lever — photograph vs watercolor vs flat vector vs 3D render are different universes.
  • 5. Lighting: the 80/20 of image quality — see the mini-bank below.
  • 6. Mood: one emotional word tunes color and atmosphere: serene, ominous, whimsical, nostalgic.
  • 7. Composition: close-up / wide shot / centered / rule of thirds / from above / empty space left for text.
  • Exclusions: what must NOT appear — full craft in Negative Prompting.

Word Banks — Fill the Slots Fast

  • Styles: photorealistic · documentary photography · cinematic still · flat vector · watercolor · oil painting · ink sketch · isometric · low-poly 3D · pixel art · paper cutout · claymation
  • Lighting: golden hour · blue hour · soft window light · dramatic side light · volumetric rays · neon glow · overcast diffused · candlelit · backlit silhouette
  • Composition: extreme close-up · medium shot · wide establishing shot · bird’s-eye view · low angle · centered symmetrical · rule of thirds · shallow depth of field · empty upper third for title
  • Mood: serene · triumphant · melancholic · cozy · ominous · playful · majestic · nostalgic
The 2–3 rule: pick two or three style/light words that agree with each other. Ten stacked descriptors don’t add detail — they fight, and the image turns to mud. Coherence beats quantity.

One Prompt, Every Tool — The Dialect Table

The description is universal; the exclusions and settings change costume per tool:

Tool family How it takes the formula Exclusions go…
ChatGPT / Gemini images Natural sentences — write the formula as flowing description; iterate conversationally (“same scene, morning light”) In the sentence: “no text, no people”
Midjourney Comma-separated phrases + parameters (–ar, –stylize) — full guide –no text, watermark
Stable Diffusion family Keyword-dense positive prompt + settings (steps, CFG) Dedicated negative field: “blurry, deformed, watermark, extra fingers”

Practical upshot: write your prompt once in formula form, then translate — sentences for chat tools, phrases+params for MJ, keywords+negative field for SD. Same image intent, three costumes.

One Subject, Five Styles — Watch Slot 4 Work

To burn in how much the style slot matters, here’s one identical scene through five media — steal whichever matches your project:

BASE SCENE (SLOTS 1–3 FIXED)“a lighthouse on a cliff, waves crashing below, stormy sky”
  • + “dramatic cinematic photography, long exposure” → movie-poster realism
  • + “flat vector illustration, limited palette of navy and cream” → modern website hero
  • + “children’s book watercolor, soft edges, gentle colors” → storybook warmth
  • + “isometric low-poly 3D render, pastel colors” → game-art charm
  • + “vintage travel poster, bold shapes, grainy texture, muted reds” → retro print

Five publishable directions, one sentence changed. When a client or a blog post needs “options,” this is how you generate them in minutes.

Recipes by What You’re Making

BLOG FEATURE IMAGE“[topic] themed abstract scene, [two brand colors], modern minimal style, high contrast, wide 16:9 composition with empty upper third for the title — no text, no watermark”
YOUTUBE THUMBNAIL“[subject] with an exaggerated surprised expression, bold saturated colors, clean simple background, high contrast, close-up, 16:9 — no text (add it yourself, always)”
SOCIAL POST (SQUARE)“[subject/scene], cohesive [palette], soft pleasing light, centered composition with breathing room, 1:1”
PRESENTATION SLIDE BG“subtle abstract [theme] background, muted [brand color] gradient, very low contrast, lots of empty space, wide — no busy details, no text”
E-COMMERCE PRODUCT“[product] on [clean surface], single soft side light, gentle shadow, minimalist commercial photography, sharp detail, square — no props, no text”
PROFILE / AVATAR“portrait illustration of [description], friendly expression, flat modern style, [color] circular background, centered, clean lines”
HERO ILLUSTRATION (WEBSITE)“[people/objects representing your service], flat vector illustration, [palette], subtle geometric background, wide with copy space on the left”
EDUCATION DIAGRAM-STYLE“clean infographic-style illustration of [concept], labeled zones left to right, flat design, white background, generous spacing — no dense text”

Iterating on Images — The Rules

  • One variable per generation. Change lighting OR style OR angle — never all three. Otherwise you can’t learn what caused what (refinement rules, visual edition).
  • Reroll with intent, not hope: if two consecutive generations disappoint the same way, the prompt is the problem — name the failure (too busy? wrong mood?) and fix that slot instead of rerolling a broken recipe
  • Anchor before exploring: get subject + composition right first in a plain style, THEN try style variations — style experiments on a broken composition waste every generation.
  • Vary the winner, not the prompt: when a result is close, use the tool’s variation feature on THAT image instead of rerolling from scratch.
  • Save winning prompts verbatim — your visual recipe book is the asset; a hit prompt with [bracketed] subject becomes a reusable template.

Universal Problems & Fixes

Problem (every tool has it) Fix
Garbled text in the image Request 1–3 words max, or generate textless and add typography in an editor — the pro route
Hands, fingers, anatomy oddities Pose hands out of focus/frame, exclude “deformed hands, extra fingers”, regenerate — improved everywhere, solved nowhere
Generic “AI-pretty” sameness Delete beautiful/stunning/4k; add specific medium + lighting + mood from the banks
Style mud (elements fighting) The 2–3 rule — cut to two style words that agree
Ignored instruction Move it earlier in the prompt; simplify the sentence; one idea per phrase
Wrong crop for the destination Set aspect/orientation BEFORE generating — composition doesn’t survive cropping

Mistakes That Waste Generations

  • Empty adjectives as strategy: beautiful, stunning, amazing, 8k, masterpiece — training saw them everywhere, so they select nothing; slots and word banks are the actual controls
  • Instruction-speak: “please create an image that shows…” — the polite wrapper is dead weight in every tool; describe the finished picture
  • Changing everything each round: new style + new light + new angle at once teaches you nothing about what worked — one variable, always
  • Generating first, choosing format later: composition dies in the crop — pick orientation for the destination before the first generation
  • Fighting text generation: spending ten rounds trying to fix garbled lettering that an editor would add perfectly in thirty seconds
  • No recipe book: every winning prompt you don’t save is a future hour you’ll spend rediscovering it — the library habit applies to images doubly

Ethics & Rights — The 60-Second Version

  • Real people: don’t generate identifiable real individuals without consent — deepfake territory, legally and ethically
  • Living artists & trademarks: describe style qualities (“bold ink lines, muted palette”) instead of prompting names/characters for commercial work
  • Commercial rights vary by tool and plan — check current terms; AI-image copyright law is still evolving and differs by country
  • Disclose where it matters: product photos that misrepresent, news-like images — transparency isn’t just ethics, it’s increasingly law

Which Tool for Which Job?

  • Quick + conversational: ChatGPT/Gemini images — iterate by chatting, no syntax to learn (ChatGPT guide · Gemini guide)
  • Aesthetic ceiling + control dials: Midjourney — the artist’s pick
  • Free, local, infinitely tweakable: Stable Diffusion family — the tinkerer’s pick, steepest curve
  • Honest answer: the formula matters more than the tool — a skilled prompt in a free tool beats a lazy prompt in the best one

Frequently Asked Questions

What is the best formula for AI image prompts?

Subject with one defining detail + action + environment + style/medium + lighting + mood + composition + exclusions. Fill 4–5 of those slots with concrete visual words and you’re ahead of 90% of prompts in any tool.

Do the same image prompts work in ChatGPT, Midjourney, and Stable Diffusion?

The descriptive core transfers completely; only the costume changes — sentences for chat tools, phrases plus –parameters for Midjourney, keywords plus a negative field for Stable Diffusion.

Why do my AI images look generic?

Empty adjectives. ‘Beautiful, stunning, high quality’ select nothing because they described everything in training. Concrete medium, lighting, and mood words are the actual controls.

How do I get text to appear correctly in AI images?

Keep it to 1–3 words at most — or generate the image clean and add typography in Canva/Photoshop. For anything professional, added text beats generated text every time.

How many style words should I use?

Two or three that agree with each other. Stacking ten descriptors makes elements fight and produces muddy results — coherence beats quantity.

How do I keep a consistent style across many images?

Fix a style phrase and palette and reuse them verbatim; use seed/reference features where the tool offers them; and change only the subject slot between generations.

Can I use AI-generated images on my blog or products?

Usually yes, subject to each tool’s plan terms — verify commercial rights, avoid trademarked content and real people, and remember copyright law for AI images is still settling and varies by country.

Which slot improves images the most for beginners?

Lighting. One specific lighting phrase (golden hour, soft window light, dramatic side light) transforms an image more than any other single addition.

Should I write image prompts in English?

For best results, yes — most models were trained predominantly on English captions, so English visual vocabulary maps most reliably. Chat-based tools can translate your idea first: describe in your language, ask for an English image prompt, then run it.

What’s the difference between this guide and the Midjourney guide?

This is the universal layer — the formula, word banks, and fixes that work in every tool. The Midjourney guide adds MJ’s specific dialect: parameters, workflow buttons, and advanced features like weights and permutations.

One formula. Every tool. 🖼️

Next: AI Video Prompting Guide — guide #25.

See the full Prompt Engineering roadmap →

AI Image Prompting Guide Infographic - Techprofree