Synthetic Data Generation

A fundamental challenge in applying supervised machine learning to scientific microscopy is the scarcity of labeled data. Pixel-level annotation (semantic segmentation masks) of microscopy images requires expert knowledge and is extremely time-consuming. Synthetic data generation offers strategies to produce large annotated datasets efficiently.

Why Synthetic Data?

Training robust segmentation models typically requires hundreds or thousands of annotated images. Several strategies exist to meet this requirement, each with different trade-offs between realism, diversity, computational cost, and annotation effort. Pick a strategy below.

JuSPICE integration

The juspice.synth_data_module module wraps classical augmentation and physics-based generation. juspice.dcgan provides the DCGAN back-end and juspice.stable_diff the Stable Diffusion back-end. All generated images should be wrapped in SPICEData and saved via save_data() to keep provenance records.

See Synthetic Data for a module-level explanation and juspice.synth_data_module for full API details.

Synthetic Data Notebooks

The following Jupyter notebooks demonstrate each synthetic data generation strategy:

See Notebook Examples for descriptions of each notebook.

References

Libraries

  • albumentations — fast, composable augmentation pipeline for images and masks

  • IOPaint — LaMa-based background-filling tool used by every augmenting generation method (Aug, PB, PB_NonGauss, SDiff, DCGAN); a separate install, not a JuSPICE dependency (see installation)

  • PyTorch — deep-learning framework underlying the DCGAN and Stable Diffusion generators

  • Hugging Face Diffusers — Stable Diffusion and other diffusion-model pipelines

See Bibliography for the publications behind the methods on this page.