Simulation List Guide

Overview

simulation-list manages the experiment registry — a SQLite database that tracks every analysis run, its parameters, and its status. The registry is used by ocean-reporting to populate the Hugo site and by update-from-remote after each sync.

The registry uses a two-step write pattern to avoid race conditions when multiple analysis scripts run in parallel:

  1. Each script writes a small staging YAML file under analyses/staging/.
  2. simulation-list merge atomically flushes all staging files into the DB.

Subcommands

# Show DB contents and pending staging files
simulation-list status [--base-dir DIR]

# Flush staging → DB
simulation-list merge  [--base-dir DIR]

# Export DB to a human-readable YAML snapshot
simulation-list to-yaml out.yaml [--base-dir DIR]

# Migrate a legacy YAML registry into the DB
simulation-list to-sqlite old_registry.yaml [--base-dir DIR]

--base-dir defaults to ./analyses.

Typical use

After a validation run, staging files accumulate in analyses/staging/. Merge them into the DB:

simulation-list status       # see what is pending
simulation-list merge        # flush into DB
simulation-list status       # confirm all merged

update-from-remote calls merge automatically as part of its pipeline.

status output

=== Simulation Registry ===
Database : ./analyses/simulation_list.db
Staging  : 3 pending records

Experiments (2):
  NS / Baseline      active   2015–2022
  NS / ObsKd         active   2015–2022

Analyses (14):
  NS / Baseline / gridded_2d_surface  …
  NS / Baseline / argo                …
  …

staging/ files

Each staging file is a small YAML named by a content hash:

analyses/staging/
    a3f7c2b1.yaml
    d8e01fa4.yaml

A typical staging file:

id: a3f7c2b1
area: NS
experiment: Baseline
analysis_type: argo
period: "2015–2022"
source: argo_ifremer
parameters: TEMP,PSAL,DOXY
n_profiles: 12458
n_obs: 487321
created: "2025-04-15T14:32:10"

The content hash ensures that re-running the same analysis produces the same file name, preventing duplicate entries in the DB.

Exporting and migrating

# Snapshot the current DB to a YAML file (for inspection or backup)
simulation-list to-yaml snapshot_2025.yaml

# Migrate from the old YAML-based registry format
simulation-list to-sqlite old_simulation_list.yaml

The to-sqlite command writes each analysis as a staging file rather than directly into the DB, so you can review the result with status before committing with merge.

Programmatic access

from lib.simulation_list_db import SimulationList

sl = SimulationList(base_dir='./analyses')

# Write a staging record from a validation script
sl.update_analysis(
    area='NS',
    experiment='Baseline',
    analysis_type='argo',
    period='2015-2022',
    n_profiles=12458,
    n_obs=487321,
)
sl.save()   # writes staging/<hash>.yaml

# Merge all pending staging records
sl.merge()

# Query the DB
experiments = sl.list_experiments()
analyses    = sl.list_analyses()