Stable Diffusion Image-to-Image

← Back to Synthetic Data Generation

Diffusion models [16] learn to reverse a gradual noising process, iteratively denoising a Gaussian noise sample into a high-fidelity image.

Image-to-image synthesis conditions the denoising on an existing input image and a text prompt:

  1. Encode the input image into a latent representation.

  2. Add a controlled amount of Gaussian noise (strength parameter).

  3. Run the denoising diffusion process conditioned on text and the noisy latent.

  4. Decode the resulting latent to image space.

This produces a new image that retains the overall structure of the input while incorporating diversity specified by the prompt and strength parameter. In the context of microscopy, this can be used to generate stylistic variants of real images.

See also