Back to Diffusers

Pipeline conventions and rules

.ai/pipelines.md

0.39.08.9 KB
Original Source

Pipeline conventions and rules

Shared reference for pipeline-related conventions, patterns, and gotchas. Linked from AGENTS.md, skills/model-integration/SKILL.md, and review-rules.md.

Prefer modular for new pipelines. Modular Diffusers is the preferred way to add a new pipeline; the standard DiffusionPipeline covered below is still supported but is no longer the default. We prefer modular especially for models that don't fit a fixed task-based structure (e.g. modality baked into the checkpoint) or that are actively evolving. The conventions below apply when you do build or review a standard pipeline.

Common pipeline conventions

When adding a new pipeline (or reviewing one), skim pipeline_flux.py, pipeline_flux2.py, pipeline_qwenimage.py, pipeline_wan.py first to establish the pattern. Most conventions (class structure, mixin set, __call__ shape — input validation → encode prompt → timesteps → latent prep → denoise loop → decode — encode_prompt / prepare_latents shape, output_type / generator / progress_bar plumbing, @torch.no_grad() on __call__, LoRA mixin, from_single_file support, etc.) are easiest to internalize by comparison rather than from a fixed list.

File structure

src/diffusers/pipelines/<model>/
  __init__.py                          # Lazy imports
  pipeline_<model>.py                  # Main pipeline (with __call__)
  pipeline_<model>_<variant>.py        # Variant pipelines (e.g. img2img, inpaint) — one file/class each
  pipeline_output.py                   # Output dataclass

Gotchas

  1. Config-derived static values: prefer __init__ attributes. Values that come from a sub-component's config (e.g. vae_scale_factor) belong as self.foo = ... in __init__ — not @property, not module-level constants. Note the getattr(...) fallback — sub-components may not be loaded when the pipeline is constructed (e.g. via from_pretrained on a partial config), so don't assume self.vae / self.transformer exists.

    python
    # don't do this — @property for static config value
    @property
    def is_turbo(self) -> bool:
        return bool(getattr(self.transformer.config, "is_turbo", False))
    
    # don't do this — module-level constant duplicating loadable config
    SAMPLE_RATE = 48000
    
    # do this — set once in __init__ with a getattr fallback (see pipeline_flux.py:209)
    def __init__(self, ..., vae, transformer, ...):
        ...
        self.register_modules(vae=vae, transformer=transformer, ...)
        self.vae_scale_factor = (
            2 ** (len(self.vae.config.block_out_channels) - 1) if getattr(self, "vae", None) else 8
        )
        self.sample_rate = int(self.vae.config.sampling_rate) if getattr(self, "vae", None) else 48000
    

    @property is reserved for per-call state — values that depend on something set inside __call__ (e.g. do_classifier_free_guidance reading self._guidance_scale).

  2. @torch.no_grad() discipline. Two failure modes:

    • Missing on __call__ entirely — causes GPU OOM from gradient accumulation during inference. Always decorate __call__ with @torch.no_grad().
    • Redundant inside helpers that __call__ already covers. The decorator puts every descendent in no-grad, so an inner with torch.no_grad(): is noise — and worse, it forecloses callers who want to invoke pipe.encode_prompt(...) with grads enabled (training, embedding optimization). Convention across diffusers (flux, qwen, flux2, stable_audio, audioldm2) is decorator-only.
  3. Reinventing logic that already exists in the repo. Check src/diffusers/guiders/ and src/diffusers/schedulers/ before adding new logic. Reuse what's already there; extend with a small kwarg for minor variations.

    • Schedulers / guiders — grep src/diffusers/guiders/ and src/diffusers/schedulers/ first. APG, CFG variants, DDIM, DPM++, flow matching Euler etc. are all already in the repo.
    • Reimplementing what the scheduler already does. Two examples below, both forms of "the scheduler should own this":
      python
      # don't do this - bypassing the scheduler entirely and rolling your own step
      for t in custom_timesteps:
          noise_pred = self.transformer(...)
          latents = latents - sigma * noise_pred   # custom Euler step, no scheduler.step()
      
      # don't do this — using the scheduler but inlining its default sigma math
      # (this is exactly what FlowMatchEulerDiscreteScheduler computes with shift=N — not a custom case)
      sigmas = np.linspace(1.0, 1.0 / num_inference_steps, num_inference_steps)
      sigmas = shift * sigmas / (1 + (shift - 1) * sigmas)
      self.scheduler.set_timesteps(sigmas=sigmas, device=device)
      
      # good — let the scheduler own it
      self.scheduler.set_timesteps(num_inference_steps=num_inference_steps, device=device)
      for t in self.scheduler.timesteps:
          noise_pred = self.transformer(...)
          latents = self.scheduler.step(noise_pred, t, latents).prev_sample
      
      If the inlined math matches the scheduler's default, walk through one row by hand to check, delete it and configure the scheduler instead.
  4. Subclassing an existing pipeline for a variant. Don't use an existing pipeline class (e.g. FluxPipeline) to override another (e.g. FluxImg2ImgPipeline) inside the core src/ codebase. Each pipeline lives in its own file with its own class, even if it shares 90% of __call__ with a sibling. Convention across diffusers — flux, sdxl, wan, qwenimage — is duplicated __call__ between img2img / text2img / inpaint variants, not subclassing. Reuse private utilities (shared schedulers, prep functions) but not the pipeline class itself.

  5. Copying a method from another pipeline without # Copied from. When you reuse a method like encode_prompt, prepare_latents, check_inputs, or _prepare_latent_image_ids from another pipeline, add a # Copied from annotation so make fix-copies keeps the two in sync. Forgetting it means future refactors to the source drift away from your copy silently — and reviewers waste time spotting near-identical code that should have been linked. The annotation grammar (decorator placement, rename syntax with with old->new, etc.) is implemented in utils/check_copies.py — read it for the exact rules.

  6. Be deliberate about methods on the pipeline. __call__ is the user's mental model. The methods on the class are how they navigate it. Diffusers convention (flux, sdxl, wan, qwenimage) is a flat class body of public lifecycle methods (__init__, check_inputs, encode_prompt, prepare_latents, __call__). Two principles, not strict rules — use judgment:

    • If a method is called from __call__, and it's a step in the pipeline lifecycle, make it public. Each call from __call__ should correspond to a step a user can identify: either a standard one (encode_prompt, prepare_latents, set_timesteps, …) or a pipeline-specific one (prepare_src_latents, prepare_reference_audio_latents, …). Don't gate these behind a _; they're part of the pipeline's API surface alongside their standard siblings.
    • If a method is only used by another method, make it private (_foo) or lift it to a module-level function — and keep the count down. Before adding one, see if the logic can be absorbed into its caller. Unless you expect the helper to be reused by another method (or another task pipeline), absorbing is usually the better call — especially when the body is small. Avoid a pipeline class littered with private helpers that bury the lifecycle..
  7. Don't modify the state of a registered component on the fly. From inside __call__ or other helper methods, don't change the state of self.text_encoder / self.transformer / self.vae — no in-place .to(dtype/device), no setting attributes/buffers or swapping submodules. Components are shared and routinely reused across pipelines, so a per-call mutation may silently change another pipeline's outputs. You should pass a component that's already in the right state, and document that expectation explicitly. Only when that's genuinely inconvenient and you must change state for the duration of a call — e.g. swapping in an attention processor — save the original first and restore it before returning, so the component is left exactly as you found it. The PAG pipelines are the reference for this: pipeline_pag_sd.py snapshots original_attn_proc = self.unet.attn_processors, installs the PAG processors for the denoising loop, then calls self.unet.set_attn_processor(original_attn_proc) at the end of __call__.

  8. Don't reimplement DiffusionPipeline. A pipeline subclass adds only pipeline-specific steps (__call__, check_inputs, encode_prompt, prepare_latents, …). Device placement, offloading, and component loading/registration already live on the base class — don't add your own; use what's there.