I/O and Tracking

The juspice.io and juspice.tracking modules are the backbone of every JuSPICE workflow. They handle loading data, saving data, and “reproducibility” — automatically keeping a record of everything you did, so someone else (or you, later) can redo it exactly. This record is kept at two levels: the whole session, and each individual dataset.


Data loading — load_data()

load_data() is the single entry point for reading supported file formats into a SPICEData container.

from juspice.io import load_data

spice = load_data("/data/image.png")

Supported formats

Extension

Back-end

.png

PIL.Image

.jpg / .jpeg

PIL.Image

.tif / .tiff

tifffile (fallback: PIL)

.spm

pySPM — AFM scan; every channel on the shared pixel grid is stacked into a (height, width, n_channels) array, with names in metadata['channel_names']

.dm3

ncempy — Gatan DigitalMicrograph TEM image; pixel spacing in metadata['pixel_size'] / ['pixel_unit']

.ser

ncempy — FEI/TIA STEM image; full acquisition parameters in metadata['acquisition']

.emi

rosettasciio — FEI/TIA project file; resolves the companion .ser data file automatically

.ibw

igor2

.jpk

nanite

.h5 / .hdf5

h5py

.npy

numpy.load

.npz

numpy.load (uses data key or first key)

This list of supported formats is continuously extended — see juspice.io for the current, authoritative set.

The returned SPICEData object contains:

  • data — the loaded numpy.ndarray,

  • metadata — source path, shape, dtype, and format metadata,

  • data_type — a string tag ("image_png", "array", etc.),

  • source_path — the absolute path of the source file,

  • history — a DatasetHistory object created automatically when track_history=True (the default).


Dataset types — single images, frame sequences, video

SPICEData natively manages three dataset shapes through its dataset_type attribute:

dataset_type

Meaning

"single_image"

Default. data is one 2D/3D array.

"multiple_frames"

An ordered image sequence. data is all frames

stacked along a new leading axis, shape (n_frames, ...).

"video"

Frames automatically extracted from a video file, internally represented the same way as "multiple_frames".

Every SPICEData exposes a uniform interface regardless of dataset_type:

spice.dataset_type   # "single_image" | "multiple_frames" | "video"
spice.n_frames       # 1 for single_image, else the frame count
spice.frames         # uniformly shaped (n_frames, ...) view of data
spice.frame_paths    # source file path(s), one per frame

Loading an ordered image-sequence folder

Plain load_data(directory) is unchanged — it still returns a SPICEDataset (one independent SPICEData per file), used for example by clustering workflows. To instead load a directory as a single ordered frame sequence, opt in explicitly:

from juspice.io import load_data, load_frame_sequence

# Explicit opt-in via load_data
spice = load_data("/data/experiment", as_frames=True)

# Or call the dedicated entry point directly
spice = load_frame_sequence("/data/experiment")

spice.dataset_type   # "multiple_frames"
spice.n_frames        # e.g. 500

Frames are ordered by the first numeric run embedded in each filename, so common conventions such as frame_0001.png, img_001.png and image_000001.tif are all detected automatically. All frames in the folder must share the same shape.

Loading a video file

Video files are detected automatically by extension (.mp4, .avi, .mov, .mkv):

from juspice.io import load_data, load_video

spice = load_data("/data/experiment.mp4")   # dispatches to load_video
# or
spice = load_video("/data/experiment.mp4")

spice.dataset_type                       # "video"
spice.metadata["fps"]                    # frames per second, if available
spice.metadata["extracted_frame_dir"]    # adjacent frames/ directory

Frames are extracted (via imageio) into a directory named after the video, created adjacent to it — experiment.mp4 produces an experiment/ folder of frame_000001.png, frame_000002.png, … On subsequent loads of the same video, an existing extraction is reused rather than re-extracted.


Data saving — save_data()

save_data() accepts a single SPICEData object and automatically determines where to write its output by inspecting the calling notebook or script filename — e.g. calling save_data(spice) from my_analysis.ipynb writes my_analysis.npy, my_analysis.json, and my_analysis_history.py next to it. No path argument is needed.

from juspice.io import save_data

save_data(spice)

Health validation

Before writing anything, save_data() validates the object via _check_spice_health(), raising ValueError if:

  • data is None or not a numpy.ndarray,

  • data is an empty array,

  • metadata is not a dict,

  • data_type or source_path is an empty string.

Auto-detected output location

The output stem and directory are inferred in this priority order:

  1. Jupyter/IPython — tries ipynbname (if installed), then the VS Code kernel variable __vsc_ipynb_file__, then falls back to the current working directory name.

  2. Plain Python script — walks the call stack to find the outermost non-JuSPICE frame and uses that file’s stem and directory.

Three artefacts are written automatically

Given a calling notebook named my_analysis.ipynb (or script my_analysis.py):

File

Content

my_analysis.npy

The full SPICEData object, serialised as a NumPy pickled object array. Reload with np.load("my_analysis.npy", allow_pickle=True)[0].

my_analysis.json

Dataset-level provenance metadata — source path, import timestamp, all recorded operations, version info.

my_analysis_history.py

Session-level history script captured by the active Tracker.

Note

save_data() calls _write_history() internally, so you do not call save() separately.

_write_history() temporarily pauses the active recording flag so that as_script() can run without raising RuntimeError. The snapshot captures all recording blocks that completed before the save_data() call — the current cell (the one containing save_data(spice)) is not yet captured at that moment, which is the desired behaviour since save_data is infrastructure, not analysis code.


Session-level tracking — Tracker

Tracker saves the exact code you wrote — word for word — across a whole notebook or script session.

Construction

from juspice.tracking import Tracker

tracker = Tracker(
    include_metadata=True,   # prepend header with date, version, git hash
    seed=42,                 # prepend numpy random seed
    notes="my analysis",     # free-text session description
)

Constructing a Tracker automatically registers it as the active tracker for the session. save_data() uses this registration to write the history file without needing the tracker passed explicitly.

Recording blocks

Code is captured between pairs of recording_start() and recording_stop() calls.

tracker.recording_start()

# --- Block 0: setup ---
from juspice.io import load_data, save_data
import numpy as np
data_path = "/data/image.png"

tracker.recording_stop()

In Jupyter notebooks, each of recording_start() and recording_stop() must occupy its own cell. The tracker uses IPython pre_execute / post_execute kernel hooks to capture full cell source after each cell completes.

In plain Python scripts, the tracker uses inspect to record the caller’s file and line number at recording_start(), then uses linecache to slice and capture the lines between start and stop when recording_stop() is called.

Control statements are filtered out

recording_start(), recording_stop(), as_script(), save(), and history are never included in the captured history.

Block separation in the history file

The generated _history.py file inserts two blank lines between successive recording blocks for readability.

State guards

Call

Raises

recording_start() while already active

RuntimeError — “Recording is already active.”

recording_stop() with no active block

RuntimeError — “Cannot stop recording …”

as_script() or save() while active

RuntimeError — “Cannot generate a history script …”

Inspecting the accumulated history

# As a string — useful for interactive inspection
print(tracker.history)

# As a standalone script string
script = tracker.as_script()

# Written to disk manually (outside any recording block)
tracker.save("analysis_history.py")

Dataset-level tracking — DatasetHistory

DatasetHistory records what happened to each SPICEData object — a structured log of operations, not the literal code that ran them. load_data() creates one automatically and attaches it as spice.history.

Automatic operation recording

Processing functions decorated with track_dataset_operation() record themselves into spice.history automatically on every call:

from juspice.tracking import track_dataset_operation

@track_dataset_operation
def invert(spice):
    spice.data = 255 - spice.data
    return spice

spice = invert(spice)               # automatically recorded
spice = invert(spice, track_history=False)  # opt-out for one call

This dataset-level history accumulates in spice.history as operations run; you do not need to save it explicitly — the next save_data() call persists it automatically to <stem>.json (see Data saving above).

Accessing the provenance

# Reconstructed, executable pipeline steps
for line in spice.history.to_lines():
    print(line)

# Get a JSON-serialisable dict (saved alongside .npy by save_data())
meta = spice.history.to_metadata_dict()

# Generate a standalone replay script
script = spice.history.to_script(output_path="result.npy")

How the two systems interact

System

Captures

Tracker

Literal code written by the analyst.

DatasetHistory

Structured operations applied to each dataset.

A single save_data() call triggers both:


See also