juspice.tracking¶
The juspice.tracking module provides two complementary reproducibility
systems that operate at different granularities:
Session-level tracking via
Tracker— captures the full sequence of code blocks executed in a notebook or script as a single ready-to-run history script.Dataset-level tracking via
DatasetHistoryandtrack_dataset_operation()— records individual operations applied to eachSPICEDataobject and generates per-dataset provenance JSON whensave_data()is called.
Both systems are independent and complementary. A typical analysis uses both.
Two-track architecture¶
┌─────────────────────────────────────────────────────────────┐
│ Analysis session │
│ │
│ Tracker (session level) │
│ ┌──────────────────────────────────────────────────────┐ │
│ │ recording_start() │ │
│ │ block 0: imports, paths, setup │ │
│ │ recording_stop() │ │
│ │ │ │
│ │ recording_start() │ │
│ │ block 1: load_data() → spice │ │
│ │ recording_stop() │ │
│ │ │ │
│ │ recording_start() │ │
│ │ block N: process(spice) │ │
│ │ save_data(spice) ← writes all 3 artefacts │ │
│ │ recording_stop() │ │
│ └──────────────────────────────────────────────────────┘ │
│ │
│ DatasetHistory (per-dataset level) │
│ ┌──────────────────────────────────────────────────────┐ │
│ │ spice = load_data(...) → spice.history created │ │
│ │ @track_dataset_operation decorates each op │ │
│ │ save_data(spice) │ │
│ │ → <stem>.npy (full SPICEData object) │ │
│ │ → <stem>.json (provenance metadata) │ │
│ │ → <stem>_history.py (session history script) │ │
│ └──────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────┘
Recording block pattern (Jupyter notebook)¶
Each recording block spans exactly one recording_start()
call, one or more content cells, and one
recording_stop() call. In Jupyter, each of
those must be in its own cell:
┌─────────────────────────────────────────────────┐
│ Cell 1 [single statement] │
│ tracker.recording_start() │
│ # This cell is NOT recorded. │
└─────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────┐
│ Cell 2 [content — freely written] │
│ # Import libraries and define paths. │
│ from juspice.io import load_data, save_data │
│ import numpy as np │
│ data_path = "/data/image.tif" │
│ # Comments are captured too. │
└─────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────┐
│ Cell 3 [single statement] │
│ tracker.recording_stop() │
│ # Cell 2 is recorded; this stop cell is NOT. │
└─────────────────────────────────────────────────┘
The recording_start() cell itself is never
recorded. The recording_stop() cell is
recorded only up to (but not including) the stop call. Any number of content
cells may appear between start and stop.
The generated _history.py file inserts two blank lines between
successive recording blocks for readability.
In plain Python scripts, recording_start()
and recording_stop() may appear anywhere in
the file on separate lines; source lines between them are read directly using
linecache and textwrap.dedent().
State guards¶
The following calls raise RuntimeError:
recording_start()when recording is already active — “Recording is already active.”recording_stop()when no recording is active — “Cannot stop recording because no recording session is active.”as_script()orsave()while recording is active — “Cannot generate a history script while recording is active.”
Note
_write_history() (called internally by
save_data()) bypasses the active-recording guard so that
history is correctly written even when save_data() is
called from inside a recording block.
Generated files¶
save_data()(save_data(spice))Determines the output stem and directory from the calling notebook or script filename automatically. Writes three artefacts with no new subfolder:
<stem>.npyThe complete
SPICEDataobject serialised as a NumPy pickled object array. Reload withnp.load("<stem>.npy", allow_pickle=True)[0].<stem>.jsonProvenance metadata in JSON format, containing: original source path, import timestamp, recorded operations with arguments and return-type metadata, serialized user-defined functions and classes, environment and version information (Python, juspice, git hash).
<stem>_history.pySession-level history script written by
_write_history()from the activeTracker. Contains the metadata header (ifinclude_metadataisTrue) followed by all recorded code blocks, separated by two blank lines.
save_script()(spice.history.save_script(path, output_path))Writes a standalone per-dataset replay script to an explicit
path. The script is structured as: Load Libraries → User Function Definitions → Main Pipeline. Call this directly when you need a per-dataset replay script at a specific path rather than the auto-named artefact fromsave_data().
Tracker¶
- class juspice.tracking.Tracker(*, include_metadata: bool = True, seed: int | None = None, notes: str = '')¶
Bases:
objectSession recorder that builds a reproducible Python script.
Create one tracker at the start of a notebook or script. Code run between
recording_start()andrecording_stop()is recorded: in Jupyter/IPython each executed cell is captured, and in a plain script the source lines between the two calls are read from the file. Get the script withas_script()or write it withsave(). The new tracker becomes the active tracker used byjuspice.io.save_data().- Parameters:
include_metadata (bool, optional) – If
True(default), start the script with comment lines giving the date, Python and juspice versions, git commit (if available), seed and notes.seed (int, optional) – If given, the script starts with
np.random.seed(seed).notes (str, optional) – Free-text description, added as a
# notes:comment.
Examples
>>> from juspice.tracking import Tracker >>> tracker = Tracker(include_metadata=True, seed=42, notes="demo run") >>> tracker.recording_start() >>> spice = juspice.io.load_data('image.tif') >>> tracker.recording_stop() >>> print(tracker.as_script())
- property history: str¶
All recorded lines joined by newlines.
- Type:
str
- recording_start() None¶
Start recording code.
In Jupyter/IPython, each cell that runs while recording is on is recorded in full. In a plain script, the file and line of this call are noted, and the lines up to
recording_stop()are recorded.- Raises:
RuntimeError – If recording is already active.
Examples
>>> tracker.recording_start() >>> spice = juspice.io.load_data('image.tif') >>> tracker.recording_stop()
- recording_stop() None¶
Stop recording code.
- Raises:
RuntimeError – If no recording block is active.
Examples
>>> tracker.recording_stop()
- as_script() str¶
Return the recorded history as a Python script.
The header comments come first, then the recorded code with at most two blank lines in a row. Calls of the deprecated
extract_features(spice, track_history=False)are rewritten to store their result in a named variable.- Returns:
Script text ending with a newline.
- Return type:
str
- Raises:
RuntimeError – If a recording block is currently active.
- save(path: str) None¶
Write the history script to a file.
- Parameters:
path (str) – Output file path, usually ending in
.py.- Raises:
RuntimeError – If a recording block is currently active.
Examples
>>> tracker.save('/tmp/analysis_script.py')
Constructor parameters
Tracker(
include_metadata: bool = True,
seed: int | None = None,
notes: str = "",
)
include_metadataWhen
True(default), a commented header block is prepended to the history script with: date/time, Python version, juspice version, git commit hash, seed, and notes.seedOptional random seed. When provided,
import numpy as npandnp.random.seed(seed)are prepended to the recorded history so that replay runs use the same seed.notesFree-text description of the analysis session stored as a
# notes:comment in the header.
Auto-registration
Constructing a Tracker automatically registers it
as the active tracker for the session. save_data() uses
this registration to write the session history file without needing the
tracker passed explicitly.
Standard notebook pattern
Every notebook should start with:
from juspice.tracking import Tracker
tracker = Tracker(include_metadata=True, notes="describe the session here")
tracker.recording_start()
Then Block 0 (all imports, path definitions, and setup code), followed by
recording_stop(). Subsequent blocks follow
the same start / content / stop pattern. Call
save_data() in the final save block — it will write the
history automatically.
Complete example
from juspice.tracking import Tracker
tracker = Tracker(include_metadata=True, seed=42, notes="demo analysis")
# Block 0 — imports and setup
tracker.recording_start()
from juspice.io import load_data, save_data, SPICEData
import numpy as np
data_path = "/data/sample.png"
tracker.recording_stop()
# Block 1 — load data
tracker.recording_start()
spice = load_data(data_path)
tracker.recording_stop()
# Block 2 — process
tracker.recording_start()
result = 255 - spice.data
spice.data = result
tracker.recording_stop()
# Block 3 — save (writes .npy + .json + _history.py automatically)
tracker.recording_start()
save_data(spice)
tracker.recording_stop()
history property
history returns the accumulated history as
a single multi-line string. Useful for inspecting recorded content
interactively before saving.
as_script()
as_script() returns the history as a
standalone Python script string.
save() is equivalent to writing this string
to disk. Both raise RuntimeError if called while a recording block
is active. Use _write_history() (called
automatically by save_data()) to bypass this guard.
DatasetHistory¶
- class juspice.tracking.DatasetHistory(source_path: str, data_type: str, import_timestamp: str, track_history: bool = True, metadata: Dict[str, ~typing.Any]=<factory>, operations: Dict[str, ~typing.Any]]=<factory>, version_info: Dict[str, str]=<factory>, definitions: Dict[str, ~typing.Dict[str, ~typing.Any]]=<factory>, callable_instances: Dict[str, ~typing.Dict[str, ~typing.Any]]=<factory>, _instance_counter: int = 0, import_function_name: str = 'load_data')¶
Bases:
objectRecord of how a dataset was loaded and processed.
One history is attached to each
SPICEDatawhen it is loaded. Results derived from a dataset usually share its history.juspice.io.save_data()writes it to a JSON file.- source_path¶
Path the data was loaded from.
- Type:
str
- data_type¶
Data type at load time.
- Type:
str
- import_timestamp¶
Load time in ISO format.
- Type:
str
- track_history¶
If
False, nothing is recorded.- Type:
bool
- metadata¶
Dataset metadata.
- Type:
dict
- operations¶
Recorded operations in order.
- Type:
list of dict
- version_info¶
Python, juspice and git versions.
- Type:
dict
- definitions¶
Source code of recorded user-defined functions and classes.
- Type:
dict
- callable_instances¶
Recorded instances of user-defined callable classes.
- Type:
dict
- import_function_name¶
Loader that created the dataset, e.g.
"load_data".- Type:
str
- classmethod from_import(*, source_path: str, data_type: str, metadata: Dict[str, Any] | None = None, track_history: bool = True, function_name: str = 'load_data') DatasetHistory¶
Create a history for newly loaded data and record the load.
- Parameters:
source_path (str) – Path the data was loaded from.
data_type (str) – Data type.
metadata (dict, optional) – Dataset metadata.
track_history (bool, optional) – If
False, nothing is recorded.function_name (str, optional) – Loader used (
"load_data","load_frame_sequence"or"load_video"), so replay scripts repeat the same call.
- Returns:
New history.
- Return type:
- record_operation(*, function_name: str, module_name: str, args: List[Any], kwargs: Dict[str, Any], returned: Any | None = None, callable_obj: Any | None = None) None¶
Add an operation to the history.
Does nothing when
track_historyisFalse.- Parameters:
function_name (str) – Operation name, e.g.
"preprocess.normalize".module_name (str) – Module that performed it.
args (list) – Positional arguments.
kwargs (dict) – Keyword arguments.
returned (object, optional) – Description of the result.
callable_obj (callable, optional) – Function that was called; user code is saved for replay.
- record_code(code: str) None¶
Add a line of code to be copied as-is into replay scripts.
- Parameters:
code (str) – Python source line.
- to_metadata_dict() Dict[str, Any]¶
Return the content of the
<stem>.jsonfile.The dict holds
source_path,data_type,import_timestamp,version_info,dataset_metadataandhistory(the code lines fromto_lines()). The rawoperationsand saved definitions are not included.- Returns:
JSON-ready record.
- Return type:
dict
- to_script(output_path: str) str¶
Build a stand-alone Python script that replays the history.
The script has an imports section, the saved user definitions, any callable instances, and a main pipeline that ends with
juspice.io.save_data(spice).- Parameters:
output_path (str) – Not used.
- Returns:
Script text.
- Return type:
str
- save_metadata(path: str) None¶
Write
to_metadata_dict()to a JSON file.- Parameters:
path (str) – Output file path.
- save_script(path: str, output_path: str) None¶
Write
to_script()to a file.- Parameters:
path (str) – Output file path.
output_path (str) – Passed to
to_script().
DatasetHistory instances are attached to
SPICEData objects returned by load_data().
They record individual operations applied to the dataset in execution order,
including:
import metadata (source path, timestamp, version info),
downstream processing operations (added automatically via
track_dataset_operation()),user-defined functions and classes serialized for replay.
You do not normally construct DatasetHistory
directly. It is created automatically by load_data() when
track_history=True (the default) and can be accessed as spice.history.
Key methods
record_operation()Records one named operation with its arguments and return-type metadata. Called automatically by
track_dataset_operation()-decorated functions.record_code()Records a raw Python code string for verbatim insertion in the replay script. Useful for recording code that does not go through a tracked function.
to_metadata_dict()Returns a JSON-serializable dict suitable for writing to
<stem>.json.to_script()Builds and returns the standalone per-dataset replay script as a string. The script is structured as: Load Libraries → User Function Definitions → Main Pipeline.
save_metadata()Writes the metadata dict to
pathas JSON.save_script()Writes the replay script to
path.
track_dataset_operation¶
- juspice.tracking.track_dataset_operation(func: Callable[[...], Any] | None = None, *, track_history: bool = True) Callable[[...], Any]¶
Record calls of a function in the dataset history.
After each call, the first argument (positional or keyword) with a
historyattribute is found, or else the return value. The call is recorded in that history, with the arguments after the first, the keyword arguments and the function’s source code. Passingtrack_history=Falseto the decorated function skips recording.Can be used as
@track_dataset_operationor@track_dataset_operation(track_history=False).- Parameters:
func (callable, optional) – Function to decorate.
track_history (bool, optional) – Default for recording when the call does not pass
track_history.
- Returns:
The wrapped function, or a decorator if func is
None.- Return type:
callable
Examples
>>> @track_dataset_operation ... def invert_image(spice_obj): ... spice_obj.data = spice_obj.data.max() - spice_obj.data ... return spice_obj
track_dataset_operation() is a decorator that wraps
any function operating on a SPICEData object so that the
call is automatically recorded in spice.history. It inspects function
arguments and return values for objects carrying a .history attribute and
calls record_operation() transparently.
Decorated functions can still accept a track_history=False keyword
argument to suppress recording for a specific call.
Example
from juspice.tracking import track_dataset_operation
@track_dataset_operation
def my_filter(spice, sigma=1.0):
"""Apply Gaussian filter to spice.data."""
from scipy.ndimage import gaussian_filter
spice.data = gaussian_filter(spice.data, sigma=sigma)
return spice
# Calling my_filter automatically records the operation in spice.history:
spice = my_filter(spice, sigma=2.0)
Module members¶
History tracking that makes JuSPICE analyses reproducible.
Two levels of history are kept:
TrackerSession level. Records the code run between
recording_start()andrecording_stop()as a ready-to-run Python script, similar toEEG.historyin EEGLAB.DatasetHistoryDataset level. Attached to each
SPICEDataand records every operation applied to it, with the values used.
track_dataset_operation() is a decorator that records calls of user
functions in the dataset history.
- class juspice.tracking.Tracker(*, include_metadata: bool = True, seed: int | None = None, notes: str = '')
Bases:
objectSession recorder that builds a reproducible Python script.
Create one tracker at the start of a notebook or script. Code run between
recording_start()andrecording_stop()is recorded: in Jupyter/IPython each executed cell is captured, and in a plain script the source lines between the two calls are read from the file. Get the script withas_script()or write it withsave(). The new tracker becomes the active tracker used byjuspice.io.save_data().- Parameters:
include_metadata (bool, optional) – If
True(default), start the script with comment lines giving the date, Python and juspice versions, git commit (if available), seed and notes.seed (int, optional) – If given, the script starts with
np.random.seed(seed).notes (str, optional) – Free-text description, added as a
# notes:comment.
Examples
>>> from juspice.tracking import Tracker >>> tracker = Tracker(include_metadata=True, seed=42, notes="demo run") >>> tracker.recording_start() >>> spice = juspice.io.load_data('image.tif') >>> tracker.recording_stop() >>> print(tracker.as_script())
- property history: str
All recorded lines joined by newlines.
- Type:
str
- recording_start() None
Start recording code.
In Jupyter/IPython, each cell that runs while recording is on is recorded in full. In a plain script, the file and line of this call are noted, and the lines up to
recording_stop()are recorded.- Raises:
RuntimeError – If recording is already active.
Examples
>>> tracker.recording_start() >>> spice = juspice.io.load_data('image.tif') >>> tracker.recording_stop()
- recording_stop() None
Stop recording code.
- Raises:
RuntimeError – If no recording block is active.
Examples
>>> tracker.recording_stop()
- as_script() str
Return the recorded history as a Python script.
The header comments come first, then the recorded code with at most two blank lines in a row. Calls of the deprecated
extract_features(spice, track_history=False)are rewritten to store their result in a named variable.- Returns:
Script text ending with a newline.
- Return type:
str
- Raises:
RuntimeError – If a recording block is currently active.
- save(path: str) None
Write the history script to a file.
- Parameters:
path (str) – Output file path, usually ending in
.py.- Raises:
RuntimeError – If a recording block is currently active.
Examples
>>> tracker.save('/tmp/analysis_script.py')
- class juspice.tracking.DatasetHistory(source_path: str, data_type: str, import_timestamp: str, track_history: bool = True, metadata: Dict[str, ~typing.Any]=<factory>, operations: Dict[str, ~typing.Any]]=<factory>, version_info: Dict[str, str]=<factory>, definitions: Dict[str, ~typing.Dict[str, ~typing.Any]]=<factory>, callable_instances: Dict[str, ~typing.Dict[str, ~typing.Any]]=<factory>, _instance_counter: int = 0, import_function_name: str = 'load_data')
Bases:
objectRecord of how a dataset was loaded and processed.
One history is attached to each
SPICEDatawhen it is loaded. Results derived from a dataset usually share its history.juspice.io.save_data()writes it to a JSON file.- source_path
Path the data was loaded from.
- Type:
str
- data_type
Data type at load time.
- Type:
str
- import_timestamp
Load time in ISO format.
- Type:
str
- track_history
If
False, nothing is recorded.- Type:
bool
- metadata
Dataset metadata.
- Type:
dict
- operations
Recorded operations in order.
- Type:
list of dict
- version_info
Python, juspice and git versions.
- Type:
dict
- definitions
Source code of recorded user-defined functions and classes.
- Type:
dict
- callable_instances
Recorded instances of user-defined callable classes.
- Type:
dict
- import_function_name
Loader that created the dataset, e.g.
"load_data".- Type:
str
- source_path: str
- data_type: str
- import_timestamp: str
- track_history: bool = True
- metadata: Dict[str, Any]
- operations: List[Dict[str, Any]]
- version_info: Dict[str, str]
- definitions: Dict[str, Dict[str, Any]]
- callable_instances: Dict[str, Dict[str, Any]]
- import_function_name: str = 'load_data'
- classmethod from_import(*, source_path: str, data_type: str, metadata: Dict[str, Any] | None = None, track_history: bool = True, function_name: str = 'load_data') DatasetHistory
Create a history for newly loaded data and record the load.
- Parameters:
source_path (str) – Path the data was loaded from.
data_type (str) – Data type.
metadata (dict, optional) – Dataset metadata.
track_history (bool, optional) – If
False, nothing is recorded.function_name (str, optional) – Loader used (
"load_data","load_frame_sequence"or"load_video"), so replay scripts repeat the same call.
- Returns:
New history.
- Return type:
- record_operation(*, function_name: str, module_name: str, args: List[Any], kwargs: Dict[str, Any], returned: Any | None = None, callable_obj: Any | None = None) None
Add an operation to the history.
Does nothing when
track_historyisFalse.- Parameters:
function_name (str) – Operation name, e.g.
"preprocess.normalize".module_name (str) – Module that performed it.
args (list) – Positional arguments.
kwargs (dict) – Keyword arguments.
returned (object, optional) – Description of the result.
callable_obj (callable, optional) – Function that was called; user code is saved for replay.
- record_code(code: str) None
Add a line of code to be copied as-is into replay scripts.
- Parameters:
code (str) – Python source line.
- to_lines() List[str]
Return the recorded operations as lines of Python code.
Calls are written with the values actually used (a call with
sigma=sigmawheresigma == 2becomessigma=2). Loads becomespice = juspice.io.load_data(...)(or the loader used), and any saves become one finaljuspice.io.save_data(spice)line. The source of a recorded user-defined function is inserted once, just before its first call.The lines start with
import juspice, plusSPICEData = juspice.io.SPICEDatawhen user functions are included andimport juspice.synth_data_modulewhen needed, so the result can run as-is. Derived results share their parent’s history, so they include the parent’s operations too.See
to_script()for a full replay script.- Returns:
Code lines in execution order, or
[]if nothing was recorded.- Return type:
list of str
Examples
>>> spice = juspice.io.load_data('image.tif') >>> spice = invert_image(spice) # @track_dataset_operation-decorated >>> juspice.io.save_data(spice) >>> spice.history.to_lines() ['import juspice', 'SPICEData = juspice.io.SPICEData', "spice = juspice.io.load_data('/abs/path/image.tif')", '', 'def invert_image(spice_obj: SPICEData) -> SPICEData:', ' ...', '', 'spice = invert_image(spice)', 'juspice.io.save_data(spice)']
- to_metadata_dict() Dict[str, Any]
Return the content of the
<stem>.jsonfile.The dict holds
source_path,data_type,import_timestamp,version_info,dataset_metadataandhistory(the code lines fromto_lines()). The rawoperationsand saved definitions are not included.- Returns:
JSON-ready record.
- Return type:
dict
- to_script(output_path: str) str
Build a stand-alone Python script that replays the history.
The script has an imports section, the saved user definitions, any callable instances, and a main pipeline that ends with
juspice.io.save_data(spice).- Parameters:
output_path (str) – Not used.
- Returns:
Script text.
- Return type:
str
- save_metadata(path: str) None
Write
to_metadata_dict()to a JSON file.- Parameters:
path (str) – Output file path.
- save_script(path: str, output_path: str) None
Write
to_script()to a file.- Parameters:
path (str) – Output file path.
output_path (str) – Passed to
to_script().
- juspice.tracking.track_dataset_operation(func: Callable[[...], Any] | None = None, *, track_history: bool = True) Callable[[...], Any]
Record calls of a function in the dataset history.
After each call, the first argument (positional or keyword) with a
historyattribute is found, or else the return value. The call is recorded in that history, with the arguments after the first, the keyword arguments and the function’s source code. Passingtrack_history=Falseto the decorated function skips recording.Can be used as
@track_dataset_operationor@track_dataset_operation(track_history=False).- Parameters:
func (callable, optional) – Function to decorate.
track_history (bool, optional) – Default for recording when the call does not pass
track_history.
- Returns:
The wrapped function, or a decorator if func is
None.- Return type:
callable
Examples
>>> @track_dataset_operation ... def invert_image(spice_obj): ... spice_obj.data = spice_obj.data.max() - spice_obj.data ... return spice_obj