Diffusion & Generative-Media Models

Parent: LLM Models and APIs · researched 2026-06-03T00:22:18.789Z· 18 sources · 12 concepts · skill diffusion-generative-media

The model family that generates continuous media (images, video, audio) by learning to reverse a noising process. A model-layer reference under the ai-agent-engineering hub (2024–2026). This is the ge

Diffusion & Generative-Media Models — Image / Video / Audio

1. The denoising-diffusion core (DDPM)

2. The score-based / SDE view

3. Latent diffusion (why almost everything runs in latent space)

4. Architectures: UNet → Diffusion Transformer

5. Flow matching & rectified flow

6. Guidance — classifier & classifier-free

7. Samplers / schedulers (the steps↔quality knob)

8. Few-step generation & distillation

9. Conditioning & control (the six levers)

10. Text-to-video & the video-diffusion stack

11. Audio, music & other modalities (brief)

12. Evaluation & efficiency

Sources

Children

Frontier under this node: Architectures: UNet → Diffusion Transformer (DiT, MMDiT/SD3, FLUX), Audio/music & other-modality diffusion (overview), Classifier-free guidance (CFG, guidance scale, negative prompts), Conditioning & control (ControlNet, T2I-Adapter, IP-Adapter, DreamBooth, LoRA, textual inversion), Denoising diffusion core (DDPM, ε/v-prediction, noise schedules), Evaluation & efficiency (FID/FVD/CLIPScore, VBench, efficiency surveys), Few-step generation & distillation (consistency models, LCM, ADD/Turbo, sCM, TurboDiffusion), Flow matching & rectified flow (conditional FM, stochastic interpolants, Diff2Flow), Latent diffusion (VAE + denoiser, Stable Diffusion), Samplers/schedulers (DDIM, DPM-Solver++, Euler/Heun, Karras sigmas), Score-based SDE/ODE view (VP/VE, probability-flow ODE, EDM/Karras), Text-to-video & video diffusion (spatiotemporal DiT, Sora/Veo/Kling/CogVideoX/Wan, temporal consistency)

← the whole tree · 3D view· how to read this page