Every prompt below has been tested on Promptchan.ai specifically. They're organized by technique: character design, scenario openers, memory anchors, image prompts, and filter softeners. Click any prompt to copy it.
Don't have Promptchan.ai yet?
Now fronts AI girlfriend chat, video & images. Still strongest as an image generator.
Affiliate link. We may earn a commission if you subscribe. How we make money.
Categories on this page: Image Prompt · 14
Image Prompt·beginner
Three-axis image prompt
[SUBJECT: who], [STYLE (what medium): photorealistic, 35mm film, anime, soft watercolor, oil painting], [LIGHT: soft morning light / golden hour / harsh flash / candlelight]. Add one mood word at the end.
Why this works
Prompt quality collapses when users stack adjectives. Forcing three distinct axes (subject / style / light) plus one mood anchor is enough to produce consistent results without overwhelming the diffusion model.
Same character as before: [2-3 DEFINING FEATURES: hair color + length, eye color, one distinguishing mark/feature]. Now in [NEW SETTING]. Keep all three features identical to prior generations.
Why this works
Diffusion models don't naturally maintain character consistency across generations. Listing 2-3 immutable features each time forces the model to anchor on them. More than 3 features and you get drift; fewer than 2 and the character morphs.
In the negative prompt field, always include: extra fingers, deformed hands, bad anatomy, duplicate, text, watermark, low quality, blurry. Product-specific additions: [ADD WHAT YOU DON'T WANT]. Don't overload: 10 items max.
Why this works
Negative prompts are more effective than positive prompts for fixing common generation failures. Hands are the #1 failure mode. The model actively needs to be told to not mess them up. Overloading negatives degrades the positive prompt's influence.
Anchor the style with one specific reference the model knows. Examples that work: "shot on Portra 400", "in the style of a Wes Anderson frame", "1990s editorial fashion photography", "studio Ghibli composition, muted palette". One reference beats five adjectives.
Why this works
Models are trained on captioned images. Specific photography / style references tokenize better than adjective stacks. One good reference phrase shifts the entire output more reliably than 'cinematic, moody, atmospheric, dramatic'.
Composition: [SPECIFIC POSE: leaning against a doorframe, one shoulder touching / sitting cross-legged on the bed, elbows on knees / standing in profile, looking back over shoulder]. Describe the pose in 6-10 words with body-part anchors, not abstract mood words.
Why this works
Diffusion models generate poses from captioned photo datasets, which describe bodies in concrete part-and-position language. 'Leaning against a doorframe' generates reliably; 'feeling relaxed' doesn't. Body-part anchors (shoulder, elbow, knee) act like keypoints the model latches onto.
Outfit is fixed: [ITEM 1: color + material], [ITEM 2: color + material], [SHOES or NO SHOES]. Do not swap items. If generating multiple images of the same character, repeat the outfit description verbatim. One detail change breaks consistency.
Why this works
Outfit drift is the second-biggest consistency failure after face drift. Listing items as color-plus-material tokens ('black cotton t-shirt') rather than style words ('casual shirt') locks the generation to specific visual features the model can reproduce across seeds.
Expression: [SPECIFIC FACIAL STATE: half-smile, eyes narrowed slightly / lips parted, brows raised / laughing with head tilted back / neutral mouth, intense eye contact]. Describe mouth and eyes separately. Avoid emotion words. Describe the face.
Why this works
Emotion words ('happy', 'sad') map to averaged expressions in training data and produce bland results. Describing mouth and eyes as independent features activates more specific regions of the model's expression space. This is how professional prompt engineers direct faces.
Environment: [ONE SPECIFIC PLACE: a kitchen with morning light through a single window / a hotel bathroom with greenish fluorescent lighting / a car interior at night, parking lot outside]. Include one light source, one surface, one background object.
Why this works
Diffusion models render environments well when given the light-source / surface / background-object triad because that matches how scene captions are typically structured in training data. Vague settings ('a cozy room') produce averaged slop; the triad produces specific, coherent scenes.
For anime output: lead with "anime style, [SUB-STYLE: 90s cel anime / modern Kyoto Animation / shounen manga / semi-realism]". Then: character features. Then: "masterpiece, best quality, detailed". Booru-style tags work better than sentences: comma-separated, no articles.
Why this works
Anime diffusion models (most Waifu Diffusion / NovelAI descendants) were trained on Danbooru/Gelbooru tags, not natural language. Booru syntax (tag-based, comma-separated, no 'a'/'the') produces sharper outputs than prose because it matches the training caption format exactly.
"Photograph of [SUBJECT], [SHOT TYPE: headshot / three-quarter / full body], shot on [CAMERA + LENS: Canon 5D with 85mm / Sony A7 with 50mm / Hasselblad], [LIGHT: natural window light / softbox key + fill]. Shallow depth of field. Skin texture visible." No painting/illustration words.
Why this works
Photo-realism degrades when you mix photography terms with art terms: the model averages between modes. Sticking strictly to camera/lens/light/depth-of-field vocabulary and explicitly including 'skin texture visible' counteracts the smoothing bias most SDXL-era models have.
Light: "golden hour, low sun from camera-left, warm rim light on hair and jawline, shadow fall-off on the right side of the face, slight haze". Don't just say 'golden hour'. Specify direction, warmth, what it touches, and what it leaves in shadow.
Why this works
'Golden hour' alone triggers a generic warm-orange preset. Specifying light direction (camera-left), affected surfaces (hair, jawline), shadow behavior, and atmosphere (haze) forces the model to render the physics rather than the stereotype. This matches how cinematographers actually describe light.
Style reference (one, not five): "in the style of [ARTIST/PHOTOGRAPHER: Helmut Newton / Nan Goldin / Peter Lindbergh / Sakimichan / Artgerm]". Only use names the model clearly knows. If unsure, substitute "in the style of [DECADE] [GENRE] [MEDIUM]", e.g. "1970s fashion editorial photography".
Why this works
Single-artist references produce stronger stylistic coherence than stacked adjectives because the model has dense clusters of captioned work per artist. Stacking multiple artists averages them into mush. The decade/genre/medium fallback works for lesser-known aesthetics and is safer than guessing at name recognition.
"Close-up portrait, face fills 60% of frame, sharp focus on eyes, shallow depth of field, background blurred to soft bokeh." Avoid full-body framing when facial consistency matters.
Why this works
Full-body shots distribute pixel budget across the whole figure, which is why faces look melted at distance. Explicit close framing with 'face fills 60% of frame' and 'sharp focus on eyes' concentrates the model's rendering attention where it matters for character identity.
"Full body shot, [SUBJECT] standing with weight on [LEFT/RIGHT] leg, [OTHER LEG] slightly bent, hands [SPECIFIC POSITION]. Floor/ground visible. Feet in frame." Specify weight-bearing leg to avoid the floating-stance failure.
Why this works
Diffusion models render full-body figures poorly because they don't naturally model weight distribution. Explicitly specifying which leg bears weight and where the other leg is positioned constrains the pose to a physically plausible one. This single instruction removes most of the uncanny-stance failures.