manage-analyses Guide

Overview

manage-analyses is the housekeeping tool for the analyses/ output tree. All subcommands default to a dry run — add --apply to actually execute.

SubcommandPurpose
cleanRemove files at the wrong directory depth (non-canonical layout)
wipeRemove all regeneratable outputs so scripts can be re-run from scratch
syncRsync the remote analyses tree to a local directory
import-statsImport existing CSV statistics files into statistics.db

clean

Removes files that do not conform to the canonical layout — for example, plots written directly under validations/<experiment>/plots/ instead of the correct plots/<domain>/<model>/<period>/<type>/ path. Canonical files are left untouched.

# Preview what would be removed
manage-analyses clean

# Remove non-conforming files
manage-analyses clean --apply

Run clean after migrating from an older layout version to sweep up any residual artefacts that fix_analyses_layout moved but did not delete.

wipe

Removes all regeneratable outputs — PNG plots and statistics files — from the entire analyses tree so that all validation scripts can be re-run cleanly.

Non-regeneratable content is never touched:

  • .md narrative files (hand-written area descriptions, experiment notes)
  • metadata.yaml
  • simulation_list.db and staging/
# Preview what would be removed
manage-analyses wipe

# Wipe all regeneratable outputs
manage-analyses wipe --apply

Regeneratable patterns removed by wipe:

*_validation_statistics.txt
*_3d_validation_statistics.txt
*_profile_validation_statistics.txt
*_tidal_statistics.txt
*_tidal_detailed_statistics.txt
*_scenario_statistics.txt
gesla_station_comparison.csv
station_comparison.csv
*_validation_statistics.csv
*_trends.txt
**/*.png

sync

Rsyncs the remote analyses tree to a local directory. Only files that conform to the canonical layout are transferred; old-layout artefacts at the wrong directory depth are automatically excluded via rsync filter rules.

# Dry-run sync (shows what would transfer — safe to run first)
manage-analyses sync kb@remote:/data/analyses/

# Actual sync
manage-analyses sync kb@remote:/data/analyses/ --apply

# Sync to a specific local directory
manage-analyses sync kb@remote:/data/analyses/ \
    --analyses-dir /data/local/analyses --apply

# Print the rsync command only (for manual inspection or tweaking)
manage-analyses sync kb@remote:/data/analyses/ --print-cmd

# Use a specific SSH key
manage-analyses sync kb@remote:/data/analyses/ --apply --ssh ~/.ssh/id_rsa

update-from-remote calls manage-analyses sync internally. Use manage-analyses sync directly when you want to sync without triggering the full report/deploy pipeline.

import-stats

Walks the analyses/ tree for all *_validation_statistics.csv files and upserts their rows into analyses/statistics.db. This is a one-off migration for results produced before the database was introduced; new runs populate the database automatically.

# Preview what would be imported (dry run)
manage-analyses import-stats

# Actually import
manage-analyses import-stats --apply

The importer infers area, experiment, domain, and model from the file’s directory path. See the Statistics Database guide for details on the schema and how to query the database.

–analyses-dir

All subcommands accept --analyses-dir DIR to set the local analyses root. There is no ./analyses fallback: when omitted, the root is expanded from OCEANICU_ANALYSES_FOLDER in this machine’s data-roots file (<hostname>_ocean-post_data_roots.yaml). If neither is set, the command stops with an error naming the missing variable — wipe in particular is destructive, so it never guesses a directory from the current path.

Canonical layout reminder

The canonical layout that clean and sync enforce (defined in lib/layout.py):

<analyses_dir>/areas/<AREA>/validations/<EXPERIMENT>/
    plots/<domain>/<model>/<period>/<type>/
    tables/<domain>/<model>/<type>/

Where:

  • <domain> is physics or bio
  • <model> is the model name, e.g. pyGETM
  • <period> is a year (2020) or a range (2015-2022)
  • <type> is surface, bottom, 3d, argo, tidal, wod, cruise, platform, ices, or a depth slice like 0050m

<EXPERIMENT> itself can be a single name (the legacy layout, still used by NS and AMM7) or <SOURCE>/<experiment> — e.g. CMEMS/tidal, CMIP6_raw/GFDL-ESM4-ssp126/run01 (NSe’s layout, selected via --source on the validation scripts). Both forms are recognised by discovery and by clean/sync; see sources.yaml and the tidal/gridded guides for how the --source form is produced.