From an Image Folder to a Time Series¶
A microscopy video is often saved as a folder of separate image files. This notebook loads such a folder as one ordered dataset and reduces each frame to a single number, turning the video into a curve that is easy to inspect.
The data are 73 frames from an in-situ transmission electron microscopy (TEM) [4] video of a SiO₂ electrode taking up lithium (lithiation).
JuSPICE Modules and Classes¶
juspice.io:load_frame_sequence,save_data,SPICEDatajuspice.clustering_module:ClusteringAccessor(viaspice.clustering)juspice.tracking:Tracker
[1]:
from juspice.tracking import Tracker
tracker = Tracker(
include_metadata=True, notes='Dataset creation from SiO2 lithiation frames'
)
tracker.recording_start()
/Users/amir/GIT_repositories/juspice_pre_release/venvs/.venv/lib/python3.12/site-packages/tqdm/auto.py:21: TqdmWarning: IProgress not found. Please update jupyter and ipywidgets. See https://ipywidgets.readthedocs.io/en/stable/user_install.html
from .autonotebook import tqdm as notebook_tqdm
[2]:
# --------- Block 0: Setup ----------
# Standard library
from pathlib import Path
import sys
import os
# Third-party
import numpy as np
import matplotlib.pyplot as plt
# JuSPICE
from juspice.io import load_frame_sequence, save_data, SPICEData
# Paths
try:
notebook_dir = Path(__file__).resolve().parent
except Exception:
notebook_dir = Path.cwd()
cur = notebook_dir
repo_root = None
for _ in range(6):
if (cur / 'juspice').exists() or (cur / 'pyproject.toml').exists():
repo_root = cur
break
if cur.parent == cur:
break
cur = cur.parent
if repo_root is None:
repo_root = notebook_dir
if str(repo_root) not in sys.path:
sys.path.insert(0, str(repo_root))
[3]:
tracker.recording_stop()
Load the frames¶
load_frame_sequence() stacks all images in the folder into one array of shape (frames, height, width). The frames are ordered by the first number in each file name.
[4]:
tracker.recording_start()
[5]:
# --------- Block 1: Load frame sequence ----------
frames_dir = os.path.join(
repo_root, 'Sample_data', 'Train_test_images', 'SiO2_lithiation', 'images', 'train'
)
# `spice` is just a variable name for the SPICEData instance load_data()
# returns here — any name would work; we use `spice` throughout these
# notebooks as an intuitive nod to JuSPICE / SPICEData.
spice = load_frame_sequence(frames_dir)
print(f'Loaded {spice.n_frames} frames from {frames_dir}')
print(
f'dataset_type: {spice.dataset_type!r}, data.shape: {spice.data.shape}, '
f'dtype: {spice.data.dtype}'
)
fig, axes = plt.subplots(1, 3, figsize=(12, 4))
for ax, idx in zip(axes, (0, spice.n_frames // 2, spice.n_frames - 1)):
ax.imshow(spice.frames[idx], cmap='gray')
ax.set_title(f'Frame {idx}')
ax.axis('off')
plt.tight_layout()
plt.show()
Loaded 73 frames from /Users/amir/GIT_repositories/juspice_pre_release/Sample_data/Train_test_images/SiO2_lithiation/images/train
dataset_type: 'multiple_frames', data.shape: (73, 256, 256), dtype: uint8
[6]:
tracker.recording_stop()
Count regions in every frame¶
Watershed segmentation [28] splits an image into regions around bright or dark centres, much like water filling separate basins. It works on one image at a time, so we loop over the frames and store the number of regions found in each.
[7]:
tracker.recording_start()
[8]:
# --------- Block 2: Watershed clustering per frame ----------
FOOTPRINT_SIZE = 15
MIN_DISTANCE = 10
n_clusters = []
for frame in spice.frames:
frame_spice = SPICEData(
data=frame, metadata={}, data_type='array', source_path=spice.source_path
)
ws_spice = frame_spice.clustering.watershed_clustering(
footprint_size=FOOTPRINT_SIZE, min_distance=MIN_DISTANCE
)
n_clusters.append(int(ws_spice.data.max()))
feature_spice = SPICEData(
data=np.array(n_clusters, dtype=np.int32),
metadata={**spice.metadata, 'feature': 'watershed_cluster_count',
'footprint_size': FOOTPRINT_SIZE, 'min_distance': MIN_DISTANCE},
data_type='array',
source_path=spice.source_path,
extra=dict(spice.extra),
history=spice.history,
)
print(f'Cluster-count time series: shape={feature_spice.data.shape}, '
f'min={feature_spice.data.min()}, max={feature_spice.data.max()}')
Cluster-count time series: shape=(73,), min=17, max=30
[9]:
tracker.recording_stop()
Plot the count over time¶
Since the frames are in recording order, the frame index acts as a time axis. The curve shows how the number of watershed regions changes during the video.
Keep in mind that this count depends on the segmentation settings (FOOTPRINT_SIZE, MIN_DISTANCE). It is an image measure, not a direct count of physical particles.
[10]:
tracker.recording_start()
[11]:
# --------- Block 3: Plot cluster-count time series ----------
time_axis = np.arange(len(feature_spice.data))
fig, ax = plt.subplots(figsize=(8, 4))
ax.plot(time_axis, feature_spice.data, 'o-')
ax.set(
xlabel='Frame index (time)',
ylabel='Number of watershed clusters',
title='SiO2 Lithiation — Watershed Cluster Count Over Time',
)
plt.tight_layout()
plt.show()
[12]:
tracker.recording_stop()
Save¶
save_data() writes the time series (.npy), its metadata (.json), and a replay script (_history.py).
[13]:
tracker.recording_start()
[14]:
# --------- Block 4: Save ----------
# Writes <stem>.npy, <stem>.json, and <stem>_history.py next to the notebook
# This folder holds several notebooks, so save_data()'s automatic
# notebook-name detection would be ambiguous when run outside a live
# Jupyter session (e.g. via nbconvert) — pass output_stem explicitly
# to always land on this notebook's own files. See
# docs/repository_structure.rst.
stem = 'io_create_dataset'
save_data(
feature_spice,
output_stem=str(notebook_dir / stem),
)
print(f'Data: {os.path.join(notebook_dir, stem + ".npy")}')
print(f'JSON sidecar: {os.path.join(notebook_dir, stem + ".json")}')
print(f'History script: {os.path.join(notebook_dir, stem + "_history.py")}')
Data: /Users/amir/GIT_repositories/juspice_pre_release/notebooks/io/io_create_dataset.npy
JSON sidecar: /Users/amir/GIT_repositories/juspice_pre_release/notebooks/io/io_create_dataset.json
History script: /Users/amir/GIT_repositories/juspice_pre_release/notebooks/io/io_create_dataset_history.py
[15]:
tracker.recording_stop()
Review the recorded steps¶
to_lines() rebuilds the recorded operations as code. The watershed loop is plain Python rather than a recorded JuSPICE call, so it does not appear here; only loading and saving do.
[16]:
tracker.recording_start()
[17]:
# --------- Block 5: Inspect human-readable history ----------
readable_lines = feature_spice.history.to_lines()
print('Reconstructed pipeline for this SPICEData object:')
print('\n'.join(readable_lines))
Reconstructed pipeline for this SPICEData object:
import juspice
spice = juspice.io.load_frame_sequence('/Users/amir/GIT_repositories/juspice_pre_release/Sample_data/Train_test_images/SiO2_lithiation/images/train')
juspice.io.save_data(spice)
[18]:
tracker.recording_stop()