juspice.io

juspice.io is the centralised I/O layer for JuSPICE. All file reading and writing in notebooks and scripts must go through the public functions defined here.

Key entry points

  • load_data() — load any supported file format (or directory) into a SPICEData container with automatic DatasetHistory initialisation. Video files load as dataset_type='video' automatically; pass as_frames=True to load a directory as dataset_type='multiple_frames' instead of a SPICEDataset.

  • load_frame_sequence() — explicit entry point for loading an ordered image-sequence folder as a single multiple_frames SPICEData.

  • load_video() — explicit entry point for loading a video file, extracting its frames into an adjacent directory.

  • save_data() — save a SPICEData object; auto-detects the output stem and directory from the calling notebook or script and writes three artefacts: <stem>.npy, <stem>.json, and <stem>_history.py.

  • load_config() — load a YAML configuration file into a plain dict.

  • save_array() — save a raw numpy.ndarray without wrapping it in SPICEData.

For a conceptual explanation of the I/O and tracking workflow see I/O and Tracking.

Loading and saving data for JuSPICE.

All file input and output goes through the functions here and the SPICEData container. A SPICEData holds one of three kinds of dataset (dataset_type): 'single_image' (default), 'multiple_frames' (an ordered image sequence) or 'video' (frames extracted from a video file).

Main functions:

class juspice.io.SPICEData(data: ndarray, metadata: dict, data_type: str, source_path: str, extra: dict = <factory>, history: DatasetHistory | None = None, dataset_type: str = 'single_image')

Bases: object

Container for any data loaded or produced by JuSPICE.

Processing interfaces are available as properties, e.g. spice.preprocess or spice.clustering.

Parameters:
  • data (np.ndarray) – The data array.

  • metadata (dict) – Information about the data, such as shape, dtype, pixel size and acquisition settings.

  • data_type (str) – Source kind, e.g. 'image_png', 'image_tif', 'image_jpeg', 'em_spm', 'em_dm3', 'em_ser', 'em_emi', 'afm', 'array', 'multiple_frames', 'video' or 'other'.

  • source_path (str) – Absolute path of the source file, or of the folder or video file for frame datasets.

  • extra (dict, optional) – Additional objects, such as raw channel data or trained models.

  • history (DatasetHistory, optional) – Record of the operations applied to this dataset.

  • dataset_type ({'single_image', 'multiple_frames', 'video'}, optional) – Frame datasets store all frames stacked along the first axis of data, shape (n_frames, ...).

data: ndarray
metadata: dict
data_type: str
source_path: str
extra: dict
history: DatasetHistory | None = None
dataset_type: str = 'single_image'
property n_frames: int

Number of frames; 1 for a single image.

Type:

int

property frames: ndarray

Data as (n_frames, ...), also for a single image.

Iterating over spice.frames always gives one frame at a time.

Type:

np.ndarray

property frame_paths: list

Source file of each frame, in frame order.

Type:

list of str

property preprocess: PreprocessingAccessor

Preprocessing methods.

Each call changes data in place, is recorded in history, and returns self:

spice.preprocess.normalize() \
     .preprocess.equalize_histogram() \
     .preprocess.binarize(threshold=0.5)

See juspice.preprocess_module.PreprocessingAccessor.

Type:

PreprocessingAccessor

property features: FeatureExtractionAccessor

Feature-extraction methods.

extract() returns the feature stack as a new SPICEData that shares this object’s history:

feature_spice = spice.features.extract(feature_names='sobel_magnitude')

See juspice.feature_extract.FeatureExtractionAccessor.

Type:

FeatureExtractionAccessor

property inpaint: InpaintingAccessor

Inpainting methods.

Each call changes data in place, is recorded in history (including the mask), and returns self:

spice.inpaint.telea(mask, radius=2.0) \
     .inpaint.biharmonic(other_mask)

See juspice.inpainting_module.InpaintingAccessor.

Type:

InpaintingAccessor

property synth_generation: SynthGenerationAccessor

Synthetic data generation.

generate() builds a new image set from a config and training folders; it does not use data. The result is a new multiple_frames SPICEData that shares this object’s history:

synth_spice = spice.synth_generation.generate(
    config=config, input_dataset="EBC1", method_name="PB",
    repo_root=repo_root, N_images=5, device="cpu",
)

See juspice.synth_data_module.SynthGenerationAccessor.

Type:

SynthGenerationAccessor

property shape_based: ShapeBasedAccessor

Drawing of synthetic shape images.

Each drawing method changes data (the canvas) in place and returns self:

spice.shape_based.create(width=256, height=256, noise_type='gaussian') \
     .shape_based.generate_gaussian_circles(num_objects=15) \
     .shape_based.generate_wiggly_cracks(num_cracks=5)

Call create() first. generate_sequence() returns a new multiple_frames SPICEData instead. See juspice.shape_based.ShapeBasedAccessor.

Type:

ShapeBasedAccessor

property clustering: ClusteringAccessor

Clustering methods.

Each method returns a new SPICEData holding the label image and sharing this object’s history; data is not changed:

ws_spice   = spice.clustering.watershed_clustering(footprint_size=60)
slic_spice = spice.clustering.slic_clustering(n_segments=200)
otsu_spice = spice.clustering.otsu_multithreshold(n_classes=2)
km_spice   = spice.clustering.kmeans_clustering(n_clusters=3)

See juspice.clustering_module.ClusteringAccessor.

Type:

ClusteringAccessor

property segmentation: UNetSegmentationAccessor

Segmentation and tracking methods.

Each method returns a new SPICEData that shares this object’s history. For example, train_unet() returns the loss per epoch:

seg_spice = spice.segmentation.train_unet(
    train_images_dir=x_train_dir,
    train_masks_dir=y_train_dir,
    arch='Unet',
    encoder_name='resnet34',
    epochs=20,
    device=device,
)

See juspice.segmentation_module.UNetSegmentationAccessor.

Type:

UNetSegmentationAccessor

class juspice.io.SPICEDataset(items: list, folder: str, metadata: dict)

Bases: object

Ordered collection of SPICEData objects from one folder.

Supports len(), indexing and iteration.

Parameters:
  • items (list of SPICEData) – Loaded files in sorted order.

  • folder (str) – Absolute path of the folder.

  • metadata (dict) – Folder information, such as the file list and count.

juspice.io.load_data(path: str, tracker: 'Tracker' | None = None, track_history: bool = True, as_frames: bool = False) → SPICEData | SPICEDataset

Load a supported file, or a folder, into a SPICEData.

The file type is chosen by the extension (not case-sensitive). A folder is loaded with load_dataset() as a SPICEDataset, or with load_frame_sequence() as one frame dataset if as_frames=True. The loaded data gets a new history, which also becomes the active one.

Parameters:
  • path (str) – File or folder to load. The absolute path is stored in SPICEData.source_path.

  • tracker (Tracker, optional) – Made the active tracker, whose session history is written by save_data().

  • track_history (bool, optional) – If False, the new history records nothing.

  • as_frames (bool, optional) – If True and path is a folder, load it as one multiple_frames SPICEData.

Returns:

Loaded data.

Return type:

SPICEData or SPICEDataset

Raises:
  • FileNotFoundError – If path does not exist.

  • ValueError – If the extension is not in the supported list.

Notes

Supported formats and readers:

  • .png, .jpg, .jpeg: Pillow

  • .tif, .tiff: tifffile, with Pillow as fallback

  • .spm (Bruker AFM): pySPM; all channels are stacked, see _load_spm()

  • .dm3 (Gatan TEM), .ser (FEI/TIA STEM): ncempy

  • .emi (FEI/TIA): rosettasciio, which reads the matching .ser

  • .ibw: igor2

  • .jpk: nanite

  • .h5, .hdf5: h5py

  • .npy, .npz: NumPy

  • .mp4, .avi, .mov, .mkv: frames extracted with imageio, see load_video()

Examples

>>> spice = load_data('/path/to/image.tif')
>>> arr = spice.data
>>> print(spice.metadata['shape'])
>>> frames = load_data('/path/to/experiment', as_frames=True)
>>> frames.n_frames
500
juspice.io.load_dataset(folder: str, tracker: 'Tracker' | None = None, track_history: bool = True, recursive: bool = False, allowed_ext: frozenset | None = None) → SPICEDataset

Load every supported file in a folder as a SPICEDataset.

Each file is loaded separately with load_data(), in sorted order.

Parameters:
  • folder (str) – Folder to load.

  • tracker (Tracker, optional) – Passed to load_data() for each file.

  • track_history (bool, optional) – If False, the new histories record nothing.

  • recursive (bool, optional) – If True, include subfolders.

  • allowed_ext (frozenset of str, optional) – Lowercase extensions to include, e.g. {'.png', '.tif'}. Defaults to all non-video formats supported by load_data().

Returns:

The loaded files and folder metadata.

Return type:

SPICEDataset

Raises:

FileNotFoundError – If folder does not exist or is not a directory.

juspice.io.load_frame_sequence(directory: str, tracker: 'Tracker' | None = None, track_history: bool = True, allowed_ext: frozenset | None = None) → SPICEData

Load a folder of ordered images as one multiple_frames SPICEData.

Frames are ordered by the first number in each file name (frame_0001.png, img_001.png, …) and stacked along the first axis.

Parameters:
  • directory (str) – Folder with the frame files.

  • tracker (Tracker, optional) – Made the active tracker, whose session history is written by save_data().

  • track_history (bool, optional) – If False, the new history records nothing.

  • allowed_ext (frozenset of str, optional) – Lowercase extensions to include. Defaults to .png, .jpg, .jpeg, .tif and .tiff.

Returns:

Frames with data of shape (n_frames, ...).

Return type:

SPICEData

Raises:
  • FileNotFoundError – If directory does not exist or is not a directory.

  • ValueError – If no supported frame files are found, or frame shapes disagree.

Examples

>>> spice = load_frame_sequence('/path/to/experiment')
>>> spice.n_frames
500
juspice.io.load_video(path: str, tracker: 'Tracker' | None = None, track_history: bool = True) → SPICEData

Load a video file as a dataset_type='video' SPICEData.

Frames are extracted with imageio into a folder next to the video with the same name (experiment.mp4 -> experiment/) and reused on later calls.

Parameters:
  • path (str) – Path to the video file.

  • tracker (Tracker, optional) – Made the active tracker, whose session history is written by save_data().

  • track_history (bool, optional) – If False, the new history records nothing.

Returns:

Frames with data of shape (n_frames, ...).

Return type:

SPICEData

Raises:
  • FileNotFoundError – If path does not exist.

  • ImportError – If imageio (with a video backend such as imageio-ffmpeg) is not installed.

juspice.io.save_data(spice: SPICEData, track_history: bool = True, output_stem: str | None = None) → None

Save a SPICEData together with its history files.

Three files are written, by default next to the calling notebook or script:

  • <stem>.npy: the whole SPICEData object (pickled);

  • <stem>.json: the dataset’s operation history and a description of the result;

  • <stem>_history.py: the session script from the active Tracker, if there is one.

<stem> is the notebook or script name without extension, or the working folder’s name if neither can be found.

Parameters:
  • spice (SPICEData) – Data to save. It is checked before anything is written.

  • track_history (bool, optional) – If True (default), record a save_data entry in the history before writing the JSON file.

  • output_stem (str or Path, optional) – Base path (with or without extension) for all three files, e.g. "notebooks/io/io_afm1". If None, it is detected with _detect_output_stem_and_dir(). Detection can fail for scripts or headless jupyter nbconvert runs in a folder with several notebooks; they would then all write to the same files. Passing output_stem avoids this.

Raises:

ValueError – If spice fails the health check.

Examples

>>> save_data(spice)
>>> save_data(spice, output_stem="notebooks/io/io_afm1")
juspice.io.load_config(path: str) → dict

Load a YAML configuration file as a dict.

An empty file gives an empty dict.

Parameters:

path (str) – Path to the .yaml or .yml configuration file.

Returns:

Parsed configuration.

Return type:

dict

Raises:
  • FileNotFoundError – If path does not exist.

  • ImportError – If PyYAML is not installed.

Examples

>>> cfg = load_config('notebooks/notebooks_parameters.yaml')
>>> print(cfg['sample_images_dir'])
juspice.io.save_array(arr: ndarray, path: str) → None

Save a plain NumPy array.

.npy files are written with numpy.save and .npz files with numpy.savez under the key data.

Parameters:
  • arr (np.ndarray) – Array to save.

  • path (str) – Output path ending in .npy or .npz.

Raises:

ValueError – If the extension is neither .npy nor .npz.

Examples

>>> save_array(my_array, '/tmp/features.npy')