Synthetic Data

juspice.synth_data_module creates extra, artificial training images for electrochemical and microscopy datasets — useful when you don’t have enough real, labeled images to train a machine-learning model. It offers several generation methods, built from simple shapes and noise or from deep-learning models.


Overview

JuSPICE covers two ways to generate synthetic images:

  1. Rule-based generation (PB, PB_NonGauss) — draws configurable particles, gradients, and noise so the result looks like a physically plausible microscopy image, without needing any training images.

  2. Deep-learning generation (DCGAN, SDiff) — trains or runs a neural network on your real images to produce new, photorealistic ones. DCGAN (juspice.dcgan) trains a small generative adversarial network from scratch on your dataset; SDiff (juspice.stable_diff) reuses a pretrained Stable Diffusion model instead.

A fifth method, Aug (plain geometric augmentation: flips, rotation, scaling, shearing), can be applied on its own or combined with any of the methods above.

Every method that augments images this way (Aug, PB, PB_NonGauss, SDiff, DCGAN) uses the external IOPaint tool afterward, to fill in the background a rotation or resize exposes — see IOPaint Installation. If a single rotation/scale/shear draw would expose too much background for IOPaint to fill convincingly, it’s automatically toned down and retried; tune how much is “too much” with the shared max_exposed_fraction setting in notebooks/notebooks_parameters.yaml.

DCGAN also supports early stopping: set early_stopping_patience (and optionally early_stopping_min_delta/early_stopping_monitor) in the config to stop training once its loss curve stops improving, instead of always running the full epochs. min_epoch_for_image sets a minimum number of epochs that must finish first, so early stopping can’t cut training short before the model has had a real chance to learn.

Native interface — generate()

spice.synth_generation is the main, recommended way to generate synthetic data — it fits the same pattern as JuSPICE’s other accessors (spice.preprocess, spice.features, spice.inpaint).

Unlike PreprocessingAccessor (which changes spice.data in place), generation produces a whole new, multi-image dataset from a YAML configuration file and a folder of training images — it has nothing to do with whatever spice.data currently holds. So generate() follows the same derived-object pattern as spice.features.extract(): the spice you call it on is left untouched and only supplies history context (for example, “generated from this reference training image”). A new SPICEData comes back instead, with dataset_type='multiple_frames', sharing that history:

from juspice.io import load_data
from juspice.synth_data_module import ConfigLoader

config = ConfigLoader("notebooks/notebooks_parameters.yaml")

# Any SPICEData works as the entry point / history anchor — typically
# one representative training image.
spice = load_data("path/to/a/training/image.png")

synth_spice = spice.synth_generation.generate(
    config=config,
    input_dataset="EBC1",
    method_name="PB",
    repo_root=repo_root,
    N_images=5,
    device="cpu",
)

synth_spice.dataset_type            # "multiple_frames"
synth_spice.n_frames                # number of generated images
synth_spice.frames                  # stacked generated images, shape (n_frames, ...)
synth_spice.metadata["masks_dir"]   # path to the corresponding generated masks

generate() sets up the output/input folders and runs generation in one call. spice.synth_generation.prepare_dataset(...) is also available for more advanced use, e.g. checking the resolved output paths before you actually generate anything.

ConfigLoader

Loads and checks generation settings from a YAML configuration file (notebooks/notebooks_parameters.yaml). See load_config().

Reproducibility

Calling generate() inside a Tracker recording block is saved word-for-word in the session’s _history.py. The returned SPICEData shares the input spice’s DatasetHistory, with one new synth_generation.generate entry recording which method and N_images you used. See it via synth_spice.history.to_lines() or in the <stem>.json sidecar that save_data() writes:

synth_spice.history.to_lines()
# ["spice = load_data('/path/to/training/image.png')",
#  "# Synthetic data: spice.synth_generation.generate(...) with method='PB', N_images=5"]

(The full config/repo_root/device arguments can’t be written as JSON, so this line describes what happened rather than being a literal, re-runnable call — the same treatment to_lines() gives any operation whose arguments can’t be faithfully reconstructed.)


See also