Lesson 00
00. Stable Diffusion Quick Start
Overview of the Stable Diffusion workflow: prompt → generate → iterate → refine with inpaint and upscale.
You will learn
- How text-to-image generation works at a high level
- The difference between base models, LoRA, and ControlNet
- A simple loop: prompt → review → adjust → regenerate
Core concepts
Stable Diffusion turns natural language into latent-space noise, then denoises it step by step into a pixel image. You control the result with four levers: prompt, model/checkpoint, sampler & steps, and seed. Changing only one lever at a time makes iteration predictable.
Base checkpoints (SDXL, SD 3.5) define overall capability. LoRA adapters nudge style or subject without replacing the whole model. ControlNet adds structural guidance from edges, depth, or pose maps.
Typical workflow
- Write a clear scene description (subject + environment + lighting).
- Pick aspect ratio and model on the generate page.
- Generate 2–4 variants; pick the best composition.
- Refine with img2img, inpaint, or upscale as needed.
Common mistakes
- Stacking too many quality tags (
8k, masterpiece, ultra detailed) without describing the scene. - Jumping to high CFG or extreme steps before fixing the prompt.
- Changing multiple variables between runs, making it hard to learn what worked.
Practice
Open the generator, write a 1–2 sentence scene description, pick 1:1, and generate. Change only one variable (lighting or style) and compare two outputs side by side.