synthlust
Home / Glossary / Image diffusion
Concept

Image diffusion

The family of AI models that generate images from text prompts. Separate from the chat LLM, and the source of most image-quality differences between apps.

When your AI companion sends you a picture, it's not generated by the chat model. A separate diffusion model, usually a variant of Stable Diffusion or a custom fine-tune, handles image generation. The chat model writes a prompt for the diffusion model to render.

This matters because image quality is largely decoupled from chat quality. An app with mediocre chat can have great images (Promptchan, Xtease), and vice versa. It's also why image quality varies wildly even within one app depending on which model is serving that month.

Good image prompts follow three axes: subject (who and what), style (what medium: photorealistic, anime, oil painting), and lighting (soft, dramatic, golden hour). Stacking adjectives beyond that degrades output quality. Negative prompts are often more effective than positive ones for fixing specific problems like malformed hands.

Prompts that use this concept

Image Prompt
Three-axis image prompt
Prompt quality collapses when users stack adjectives. Forcing three distinct axes (subject / style / light) plus one mood anchor is enough to produce consistent results without overwhelming the diffusion model.
Image Prompt
Character-consistency prompt
Diffusion models don't naturally maintain character consistency across generations. Listing 2-3 immutable features each time forces the model to anchor on them. More than 3 features and you get drift; fewer than 2 and the character morphs.
Image Prompt
Style anchor phrase
Models are trained on captioned images. Specific photography / style references tokenize better than adjective stacks. One good reference phrase shifts the entire output more reliably than 'cinematic, moody, atmospheric, dramatic'.
Image Prompt
Specific-pose composition
Diffusion models generate poses from captioned photo datasets, which describe bodies in concrete part-and-position language. 'Leaning against a doorframe' generates reliably; 'feeling relaxed' doesn't. Body-part anchors (shoulder, elbow, knee) act like keypoints the model latches onto.
Image Prompt
Outfit-consistency lock
Outfit drift is the second-biggest consistency failure after face drift. Listing items as color-plus-material tokens ('black cotton t-shirt') rather than style words ('casual shirt') locks the generation to specific visual features the model can reproduce across seeds.
Image Prompt
Facial-expression specificity
Emotion words ('happy', 'sad') map to averaged expressions in training data and produce bland results. Describing mouth and eyes as independent features activates more specific regions of the model's expression space. This is how professional prompt engineers direct faces.
Image Prompt
Environment grounding
Diffusion models render environments well when given the light-source / surface / background-object triad because that matches how scene captions are typically structured in training data. Vague settings ('a cozy room') produce averaged slop; the triad produces specific, coherent scenes.
Image Prompt
Anime-style token stack
Anime diffusion models (most Waifu Diffusion / NovelAI descendants) were trained on Danbooru/Gelbooru tags, not natural language. Booru syntax (tag-based, comma-separated, no 'a'/'the') produces sharper outputs than prose because it matches the training caption format exactly.
Image Prompt
Photorealistic portrait stack
Photo-realism degrades when you mix photography terms with art terms: the model averages between modes. Sticking strictly to camera/lens/light/depth-of-field vocabulary and explicitly including 'skin texture visible' counteracts the smoothing bias most SDXL-era models have.
Image Prompt
Golden-hour light recipe
'Golden hour' alone triggers a generic warm-orange preset. Specifying light direction (camera-left), affected surfaces (hair, jawline), shadow behavior, and atmosphere (haze) forces the model to render the physics rather than the stereotype. This matches how cinematographers actually describe light.
Image Prompt
Artist/photographer reference phrase
Single-artist references produce stronger stylistic coherence than stacked adjectives because the model has dense clusters of captioned work per artist. Stacking multiple artists averages them into mush. The decade/genre/medium fallback works for lesser-known aesthetics and is safer than guessing at name recognition.
Image Prompt
Face-focused framing
Full-body shots distribute pixel budget across the whole figure, which is why faces look melted at distance. Explicit close framing with 'face fills 60% of frame' and 'sharp focus on eyes' concentrates the model's rendering attention where it matters for character identity.
Image Prompt
Full-body with weight-bearing detail
Diffusion models render full-body figures poorly because they don't naturally model weight distribution. Explicitly specifying which leg bears weight and where the other leg is positioned constrains the pose to a physically plausible one. This single instruction removes most of the uncanny-stance failures.
Related concepts
Put this into practice

Our prompt library shows these techniques in real, copy-ready prompts, tested across 22 AI companion apps.

Browse prompts →