Synthetic Training Images from Particle Models¶
Instead of transforming labelled images, we can build new ones: remove the objects from a real image, then draw model particles back in. Because we place every particle ourselves, the mask is known exactly and no manual labelling is needed.
This notebook compares two particle models on the EBC1 SEM dataset: PB (round particles with a Gaussian brightness profile) and PB_NonGauss (particles with irregular outlines).
JuSPICE Modules and Classes¶
juspice.io:SPICEData,load_data,save_datajuspice.synth_data_module:ConfigLoader,SynthGenerationAccessor(viaspice.synth_generation)juspice.tracking:Tracker
[1]:
from juspice.tracking import Tracker
tracker = Tracker(
include_metadata=True, notes='Physics-based synthetic data generation'
)
tracker.recording_start()
/Users/amir/GIT_repositories/juspice_pre_release/venvs/.venv/lib/python3.12/site-packages/tqdm/auto.py:21: TqdmWarning: IProgress not found. Please update jupyter and ipywidgets. See https://ipywidgets.readthedocs.io/en/stable/user_install.html
from .autonotebook import tqdm as notebook_tqdm
[2]:
# --------- Block 0: Setup ----------
# Standard library
from pathlib import Path
import sys
import os
import json
import logging
import shutil
# Third-party
import numpy as np
import matplotlib.pyplot as plt
try:
import torch
_torch_available = True
except ImportError:
torch = None
_torch_available = False
print("WARNING: torch is not installed. Device selection will default to 'cpu'.")
# JuSPICE
from juspice.io import SPICEData, load_data, save_data
from juspice.synth_data_module import ConfigLoader
# Paths
try:
notebook_dir = Path(__file__).resolve().parent
except Exception:
notebook_dir = Path.cwd()
cur = notebook_dir
repo_root = None
for _ in range(6):
if (cur / 'juspice').exists() or (cur / 'pyproject.toml').exists():
repo_root = cur
break
if cur.parent == cur:
break
cur = cur.parent
if repo_root is None:
repo_root = notebook_dir
if str(repo_root) not in sys.path:
sys.path.insert(0, str(repo_root))
# Logging configuration
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s - %(levelname)s - %(message)s',
datefmt='%Y-%m-%d %H:%M:%S',
)
[3]:
tracker.recording_stop()
Settings¶
We choose the fastest available device and load the generation parameters from notebooks_parameters.yaml.
[4]:
tracker.recording_start()
[5]:
# --------- Block 1: Config and device ----------
_mps_available = (
torch is not None
and hasattr(torch.backends, 'mps')
and torch.backends.mps.is_available()
)
if torch is not None and torch.cuda.is_available():
device = 'cuda'
elif _mps_available:
device = 'mps'
else:
device = 'cpu'
logging.info(f'Using device: {device}.')
config_path = os.path.join(repo_root, 'notebooks', 'notebooks_parameters.yaml')
config = ConfigLoader(config_path)
logging.info(f'Configuration loaded from: {config_path}')
dataset = 'EBC1'
20:08:2026 16:54:54 - Using device: mps.
20:08:2026 16:54:54 - Configuration loaded from: /Users/amir/GIT_repositories/juspice_pre_release/notebooks/notebooks_parameters.yaml
[6]:
tracker.recording_stop()
The particle models¶
Background. The original objects are removed from real images by inpainting (Telea’s method [48]). The remaining background keeps the real noise, texture, and uneven illumination.
``PB``: Gaussian particles. Each particle is a spot whose brightness falls off like a bell curve (Gaussian) from its centre. This mimics how an imaging system blurs a small object into a soft spot; in optical microscopy this blur is called the point spread function. For SEM images it is a simplifying assumption rather than an exact physical model.
``PB_NonGauss``: irregular particles. Wavy (sinusoidal) distortions of the outline make the particles less perfectly round, and a short frame sequence shows them changing gradually.
Parameter |
Controls |
|---|---|
|
width of the Gaussian brightness profile (pixels) |
|
particle size range and number of particles |
|
how strongly overlapping particles add up |
|
how smoothly particles merge into the background |
|
extra brightness at the particle centre |
|
frequency and size of the two outline wobbles |
|
reference particle scale and growth per frame |
We first load one training image as the starting point, then generate images with PB.
[7]:
tracker.recording_start()
[8]:
# --------- Block 2: PB (Gaussian-particle) generation ----------
method_name = 'PB'
# Load one training image as the SPICEData entry point / history anchor —
# spice.synth_generation.generate() doesn't transform spice.data itself
# (it produces a brand-new dataset), but the returned dataset shares this
# object's history lineage, so the provenance reads "generated from this
# reference training set".
dataset_params = config.get_dataset_params(dataset)
training_images_dir = os.path.join(repo_root, dataset_params['training_images_dir'])
_img_exts = ('.png', '.jpg', '.jpeg', '.tif', '.tiff')
_seed_candidates = sorted(
f for f in os.listdir(training_images_dir) if f.lower().endswith(_img_exts)
)
seed_image_name = _seed_candidates[0]
# `spice` is just a variable name for the SPICEData instance load_data()
# returns here — any name would work; we use `spice` throughout these
# notebooks as an intuitive nod to JuSPICE / SPICEData.
spice = load_data(os.path.join(training_images_dir, seed_image_name))
N_images = 5
print(f'Generating synthetic images using {method_name}...')
synth_spice_pb = spice.synth_generation.generate(
config=config,
input_dataset=dataset,
method_name=method_name,
repo_root=repo_root,
N_images=N_images,
device=device,
)
print(f'Generation complete! dataset_type={synth_spice_pb.dataset_type!r}, '
f'n_frames={synth_spice_pb.n_frames}')
20:08:2026 16:54:54 - Parameters for the model were set.
20:08:2026 16:54:54 - Preprocessing training images...
Generating synthetic images using PB...
Preprocessing images: 100%|██████████| 5/5 [00:00<00:00, 62.57it/s]
20:08:2026 16:54:54 - Training images and masks were preprocessed.
20:08:2026 16:54:54 - Test images and masks were preprocessed.
20:08:2026 16:54:54 - *******************
20:08:2026 16:54:54 - Inpainting the backgrounds of the test images.
Preprocessing complete.
Images saved to: /Users/amir/GIT_repositories/juspice_pre_release/preprocessed_data/PB_EBC1/input_images
Masks saved to: /Users/amir/GIT_repositories/juspice_pre_release/preprocessed_data/PB_EBC1/input_masks
Generating synthetic images for image 0: 100%|██████████| 1/1 [00:07<00:00, 7.10s/it]
Generating synthetic images for image 1: 100%|██████████| 1/1 [00:05<00:00, 5.63s/it]
Generating synthetic images for image 2: 100%|██████████| 1/1 [00:05<00:00, 5.86s/it]
Generating synthetic images for image 3: 100%|██████████| 1/1 [00:05<00:00, 5.63s/it]
Generating synthetic images for image 4: 100%|██████████| 1/1 [00:05<00:00, 5.61s/it]
20:08:2026 16:55:25 - Synthetic image and masks were generated and saved.
Generation complete! dataset_type='multiple_frames', n_frames=15
[9]:
tracker.recording_stop()
PB results¶
Each synthetic image (top) comes with its exact mask (bottom).
[10]:
tracker.recording_start()
[11]:
# --------- Block 3: Visualize PB results ----------
output_masks_dir = synth_spice_pb.metadata['masks_dir']
n_show = min(4, synth_spice_pb.n_frames)
fig, axes = plt.subplots(2, n_show, figsize=(16, 8))
fig.suptitle('Physics-Based Generation (PB)', fontsize=16, fontweight='bold')
for idx in range(n_show):
img_data = synth_spice_pb.frames[idx]
frame_name = os.path.basename(synth_spice_pb.frame_paths[idx])
mask_path = os.path.join(output_masks_dir, frame_name)
axes[0, idx].imshow(img_data, cmap='gray')
axes[0, idx].axis('off')
if os.path.exists(mask_path):
mask_data = load_data(mask_path).data
axes[1, idx].imshow(mask_data, cmap='gray')
axes[1, idx].axis('off')
plt.tight_layout()
plt.show()
[12]:
tracker.recording_stop()
PB_NonGauss results¶
The same procedure with irregular particle outlines. Compare their shapes with the round PB particles above.
[13]:
tracker.recording_start()
[14]:
# --------- Block 4: PB_NonGauss generation ----------
method_name_ng = 'PB_NonGauss'
N_images_ng = 1
print(f'Generating with {method_name_ng}...')
# Reusing the same anchor `spice` chains this generation onto the same
# shared history as the PB run above (spice.synth_generation.generate()
# never mutates spice itself).
synth_spice_pb2 = spice.synth_generation.generate(
config=config,
input_dataset=dataset,
method_name=method_name_ng,
repo_root=repo_root,
N_images=N_images_ng,
device=device,
)
print(f'Generation complete! dataset_type={synth_spice_pb2.dataset_type!r}, '
f'n_frames={synth_spice_pb2.n_frames}')
output_masks_dir2 = synth_spice_pb2.metadata['masks_dir']
n_show2 = synth_spice_pb2.n_frames
fig, axes = plt.subplots(2, max(1, n_show2), figsize=(16, 8))
if n_show2 == 1:
axes = axes.reshape(2, 1)
fig.suptitle(
'Non-Gaussian Physics-Based Generation (PB_NonGauss)',
fontsize=16, fontweight='bold',
)
for idx in range(n_show2):
img_data = synth_spice_pb2.frames[idx]
frame_name2 = os.path.basename(synth_spice_pb2.frame_paths[idx])
mask_path = os.path.join(output_masks_dir2, frame_name2)
axes[0, idx].imshow(img_data, cmap='gray')
axes[0, idx].axis('off')
if os.path.exists(mask_path):
mask_data = load_data(mask_path).data
axes[1, idx].imshow(mask_data, cmap='gray')
axes[1, idx].axis('off')
plt.tight_layout()
plt.show()
20:08:2026 16:55:25 - Parameters for the model were set.
20:08:2026 16:55:25 - Preprocessing training images...
Generating with PB_NonGauss...
Preprocessing images: 100%|██████████| 1/1 [00:00<00:00, 39.19it/s]
20:08:2026 16:55:25 - Training images and masks were preprocessed.
20:08:2026 16:55:25 - Test images and masks were preprocessed.
20:08:2026 16:55:25 - *******************
20:08:2026 16:55:25 - Inpainting the backgrounds of the test images.
Preprocessing complete.
Images saved to: /Users/amir/GIT_repositories/juspice_pre_release/preprocessed_data/PB_NonGauss_EBC1/input_images
Masks saved to: /Users/amir/GIT_repositories/juspice_pre_release/preprocessed_data/PB_NonGauss_EBC1/input_masks
**** Frame 1/5 generated and added to the sequence. ****
**** Frame 2/5 generated and added to the sequence. ****
**** Frame 3/5 generated and added to the sequence. ****
**** Frame 4/5 generated and added to the sequence. ****
**** Frame 5/5 generated and added to the sequence. ****
Saving frames: 100%|██████████| 5/5 [00:00<00:00, 84.19it/s]
Saving frames: 100%|██████████| 5/5 [00:00<00:00, 87.95it/s]
Generating synthetic images for frame 0: 0%| | 0/5 [00:00<?, ?it/s]
**** Frame 1/1 generated and added to the sequence. ****
Generating synthetic images for frame 0: 20%|██ | 1/5 [00:05<00:23, 5.79s/it]
**** Frame 1/1 generated and added to the sequence. ****
**** Frame 1/1 generated and added to the sequence. ****
Generating synthetic images for frame 0: 40%|████ | 2/5 [00:11<00:17, 5.77s/it]
**** Frame 1/1 generated and added to the sequence. ****
**** Frame 1/1 generated and added to the sequence. ****
Generating synthetic images for frame 0: 60%|██████ | 3/5 [00:17<00:11, 5.75s/it]
**** Frame 1/1 generated and added to the sequence. ****
**** Frame 1/1 generated and added to the sequence. ****
Generating synthetic images for frame 0: 80%|████████ | 4/5 [00:23<00:05, 5.74s/it]
**** Frame 1/1 generated and added to the sequence. ****
**** Frame 1/1 generated and added to the sequence. ****
Generating synthetic images for frame 0: 100%|██████████| 5/5 [00:28<00:00, 5.78s/it]
20:08:2026 16:55:55 - =================================================
20:08:2026 16:55:55 - Synthetic image and masks were generated and saved.
**** Frame 1/1 generated and added to the sequence. ****
Generation complete! dataset_type='multiple_frames', n_frames=7
[15]:
tracker.recording_stop()
Save and review¶
We save the PB_NonGauss result. Its history contains both generation calls, because both started from the same loaded image.
[16]:
tracker.recording_start()
[17]:
# --------- Block 5: Save ----------
if synth_spice_pb2.n_frames == 0:
raise RuntimeError('No generated images found to export.')
# Writes <stem>.npy, <stem>.json, and <stem>_history.py next to the notebook
# This folder holds several notebooks, so save_data()'s automatic
# notebook-name detection would be ambiguous when run outside a live
# Jupyter session (e.g. via nbconvert) — pass output_stem explicitly
# to always land on this notebook's own files. See
# docs/repository_structure.rst.
stem = 'example_PB'
save_data(
synth_spice_pb2,
output_stem=str(notebook_dir / stem),
)
print(f'Data: {os.path.join(notebook_dir, stem + ".npy")}')
print(f'JSON sidecar: {os.path.join(notebook_dir, stem + ".json")}')
print(f'History script: {os.path.join(notebook_dir, stem + "_history.py")}')
Data: /Users/amir/GIT_repositories/juspice_pre_release/notebooks/synthetic_data/example_PB.npy
JSON sidecar: /Users/amir/GIT_repositories/juspice_pre_release/notebooks/synthetic_data/example_PB.json
History script: /Users/amir/GIT_repositories/juspice_pre_release/notebooks/synthetic_data/example_PB_history.py
[18]:
# --------- Block 6: Validate + Inspect human-readable history ----------
# synth_spice_pb2.history.to_lines() reconstructs the full synthesis pipeline
# applied to this SPICEData object into readable, runnable code: the initial
# load_data(...) call followed by each synth_generation.generate() call, and
# a trailing save_data(synth_spice_pb2).
readable_lines = synth_spice_pb2.history.to_lines()
print('Reconstructed synthesis pipeline for this SPICEData object:')
print('\n'.join(readable_lines[:6]), '...')
Reconstructed synthesis pipeline for this SPICEData object:
import juspice
import juspice.synth_data_module
spice = juspice.io.load_data('/Users/amir/GIT_repositories/juspice_pre_release/Sample_data/Train_test_images/EBC1/train/010417#4_S2480009.tif')
synth_spice = spice.synth_generation.generate(config=juspice.synth_data_module.ConfigLoader('/Users/amir/GIT_repositories/juspice_pre_release/notebooks/notebooks_parameters.yaml'), input_dataset='EBC1', method_name='PB', repo_root='/Users/amir/GIT_repositories/juspice_pre_release', N_images=5, device='mps')
synth_spice = spice.synth_generation.generate(config=juspice.synth_data_module.ConfigLoader('/Users/amir/GIT_repositories/juspice_pre_release/notebooks/notebooks_parameters.yaml'), input_dataset='EBC1', method_name='PB_NonGauss', repo_root='/Users/amir/GIT_repositories/juspice_pre_release', N_images=1, device='mps')
juspice.io.save_data(spice) ...
[19]:
tracker.recording_stop()