Illustration generated with AI; not experimental data.¶
I/O and Tracking¶
The juspice.io and juspice.tracking modules are the backbone
of every JuSPICE workflow. They handle loading data, saving data, and
“reproducibility” — automatically keeping a record of everything you did,
so someone else (or you, later) can redo it exactly. This record is kept
at two levels: the whole session, and each individual dataset.
Data loading — load_data()¶
load_data() is the single entry point for reading supported
file formats into a SPICEData container.
from juspice.io import load_data
spice = load_data("/data/image.png")
Supported formats
Extension |
Back-end |
|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
This list of supported formats is continuously extended — see
juspice.io for the current, authoritative set.
The returned SPICEData object contains:
data— the loadednumpy.ndarray,metadata— source path, shape, dtype, and format metadata,data_type— a string tag ("image_png","array", etc.),source_path— the absolute path of the source file,history— aDatasetHistoryobject created automatically whentrack_history=True(the default).
Dataset types — single images, frame sequences, video¶
SPICEData natively manages three dataset shapes through
its dataset_type attribute:
|
Meaning |
|---|---|
|
Default. |
|
|
|
Frames automatically extracted from a video file,
internally represented the same way as
|
Every SPICEData exposes a uniform interface regardless
of dataset_type:
spice.dataset_type # "single_image" | "multiple_frames" | "video"
spice.n_frames # 1 for single_image, else the frame count
spice.frames # uniformly shaped (n_frames, ...) view of data
spice.frame_paths # source file path(s), one per frame
Loading an ordered image-sequence folder
Plain load_data(directory) is unchanged — it still returns a
SPICEDataset (one independent
SPICEData per file), used for example by clustering
workflows. To instead load a directory as a single ordered frame sequence,
opt in explicitly:
from juspice.io import load_data, load_frame_sequence
# Explicit opt-in via load_data
spice = load_data("/data/experiment", as_frames=True)
# Or call the dedicated entry point directly
spice = load_frame_sequence("/data/experiment")
spice.dataset_type # "multiple_frames"
spice.n_frames # e.g. 500
Frames are ordered by the first numeric run embedded in each filename, so
common conventions such as frame_0001.png, img_001.png and
image_000001.tif are all detected automatically. All frames in the
folder must share the same shape.
Loading a video file
Video files are detected automatically by extension (.mp4, .avi,
.mov, .mkv):
from juspice.io import load_data, load_video
spice = load_data("/data/experiment.mp4") # dispatches to load_video
# or
spice = load_video("/data/experiment.mp4")
spice.dataset_type # "video"
spice.metadata["fps"] # frames per second, if available
spice.metadata["extracted_frame_dir"] # adjacent frames/ directory
Frames are extracted (via imageio) into a directory named after the
video, created adjacent to it — experiment.mp4 produces an
experiment/ folder of frame_000001.png, frame_000002.png, …
On subsequent loads of the same video, an existing extraction is reused
rather than re-extracted.
Data saving — save_data()¶
save_data() accepts a single SPICEData
object and automatically determines where to write its output by inspecting the
calling notebook or script filename — e.g. calling save_data(spice) from
my_analysis.ipynb writes my_analysis.npy, my_analysis.json, and
my_analysis_history.py next to it. No path argument is needed.
from juspice.io import save_data
save_data(spice)
Health validation
Before writing anything, save_data() validates the object
via _check_spice_health(), raising ValueError if:
dataisNoneor not anumpy.ndarray,datais an empty array,metadatais not adict,data_typeorsource_pathis an empty string.
Auto-detected output location
The output stem and directory are inferred in this priority order:
Jupyter/IPython — tries
ipynbname(if installed), then the VS Code kernel variable__vsc_ipynb_file__, then falls back to the current working directory name.Plain Python script — walks the call stack to find the outermost non-JuSPICE frame and uses that file’s stem and directory.
Three artefacts are written automatically
Given a calling notebook named my_analysis.ipynb (or script my_analysis.py):
File |
Content |
|---|---|
|
The full |
|
Dataset-level provenance metadata — source path, import timestamp, all recorded operations, version info. |
|
Session-level history script captured by the active
|
Note
save_data() calls _write_history()
internally, so you do not call save()
separately.
_write_history() temporarily pauses the
active recording flag so that as_script()
can run without raising RuntimeError. The snapshot captures all
recording blocks that completed before the save_data()
call — the current cell (the one containing save_data(spice)) is not
yet captured at that moment, which is the desired behaviour since
save_data is infrastructure, not analysis code.
Session-level tracking — Tracker¶
Tracker saves the exact code you wrote — word
for word — across a whole notebook or script session.
Construction
from juspice.tracking import Tracker
tracker = Tracker(
include_metadata=True, # prepend header with date, version, git hash
seed=42, # prepend numpy random seed
notes="my analysis", # free-text session description
)
Constructing a Tracker automatically registers it
as the active tracker for the session. save_data() uses
this registration to write the history file without needing the tracker passed
explicitly.
Recording blocks
Code is captured between pairs of recording_start()
and recording_stop() calls.
tracker.recording_start()
# --- Block 0: setup ---
from juspice.io import load_data, save_data
import numpy as np
data_path = "/data/image.png"
tracker.recording_stop()
In Jupyter notebooks, each of recording_start()
and recording_stop() must occupy its own cell.
The tracker uses IPython pre_execute / post_execute kernel hooks to
capture full cell source after each cell completes.
In plain Python scripts, the tracker uses inspect to record the
caller’s file and line number at recording_start(),
then uses linecache to slice and capture the lines between start and
stop when recording_stop() is called.
Control statements are filtered out
recording_start(),
recording_stop(),
as_script(),
save(), and
history are never included in the captured
history.
Block separation in the history file
The generated _history.py file inserts two blank lines between successive
recording blocks for readability.
State guards
Call |
Raises |
|---|---|
|
|
|
|
|
|
Inspecting the accumulated history
# As a string — useful for interactive inspection
print(tracker.history)
# As a standalone script string
script = tracker.as_script()
# Written to disk manually (outside any recording block)
tracker.save("analysis_history.py")
Dataset-level tracking — DatasetHistory¶
DatasetHistory records what happened to each
SPICEData object — a structured log of operations,
not the literal code that ran them. load_data() creates
one automatically and attaches it as spice.history.
Automatic operation recording
Processing functions decorated with track_dataset_operation()
record themselves into spice.history automatically on every call:
from juspice.tracking import track_dataset_operation
@track_dataset_operation
def invert(spice):
spice.data = 255 - spice.data
return spice
spice = invert(spice) # automatically recorded
spice = invert(spice, track_history=False) # opt-out for one call
This dataset-level history accumulates in spice.history as operations run;
you do not need to save it explicitly — the next
save_data() call persists it automatically to
<stem>.json (see Data saving above).
Accessing the provenance
# Reconstructed, executable pipeline steps
for line in spice.history.to_lines():
print(line)
# Get a JSON-serialisable dict (saved alongside .npy by save_data())
meta = spice.history.to_metadata_dict()
# Generate a standalone replay script
script = spice.history.to_script(output_path="result.npy")
How the two systems interact¶
System |
Captures |
|---|---|
Literal code written by the analyst. |
|
Structured operations applied to each dataset. |
A single save_data() call triggers both:
DatasetHistoryserialises to<stem>.json,The active
Trackerserialises to<stem>_history.py.
See also
juspice.io — full API reference for
juspice.iojuspice.tracking — full API reference for
juspice.tracking