Tracker architecture

← Back to Developer Guide

juspice.tracking exposes two independent tracking systems:

Session-level tracking (Tracker)

Tracker captures the literal code written by the analyst across an entire notebook or script session. Recording is controlled explicitly with recording_start() and recording_stop() block delimiters.

  • Constructing a Tracker automatically registers it as the active tracker via _set_active_tracker(self) so that save_data() can write the session history without receiving the tracker explicitly.

  • In Jupyter/IPython, the tracker registers pre_execute and post_execute hooks via the IPython event system. The cell containing recording_start() is never recorded. The cell containing recording_stop() causes an extra blank line to be appended so the generated history file has two blank lines between successive blocks.

  • In plain Python scripts, recording_start() uses inspect to note the caller’s file and line number; recording_stop() uses linecache and textwrap.dedent() to read the source lines between start and stop from disk.

  • Tracker-control statements (recording_start(), recording_stop(), as_script(), save(), history) are filtered from captured content.

  • State is guarded: calling recording_start() twice, calling recording_stop() without a prior start, or calling save() / as_script() while recording is active each raise RuntimeError with a descriptive message.

  • _write_history() bypasses the active-recording guard so that save_data() can write the history even when called from inside a recording block.

Dataset-level tracking (DatasetHistory + track_dataset_operation())

DatasetHistory is attached to SPICEData objects by load_data(). The track_dataset_operation() decorator wraps processing functions so that every call is recorded in spice.history automatically. When save_data() is called, the accumulated history is serialized to a <stem>.json sidecar as a history list of human-readable, executable pipeline steps.

User-defined functions passed as arguments to tracked operations are serialized by capturing their source code via inspect.getsource(). Dependency functions and classes are discovered recursively from closure variables and registered in definitions in topological order so that the per-dataset replay script is self-contained.

save_data — auto-path artefact generation

save_data() accepts only a SPICEData object. It determines the output stem and directory by inspecting the calling notebook or script filename, validates the object via _check_spice_health(), then writes:

  • <stem>.npy — full SPICEData object (pickled NumPy object array),

  • <stem>.json — dataset-level provenance JSON,

  • <stem>_history.py — session-level history from the active Tracker.