manage-analyses Guide
Overview
manage-analyses is the housekeeping tool for the analyses/ output tree.
All subcommands default to a dry run — add --apply to actually execute.
| Subcommand | Purpose |
|---|---|
clean | Remove files at the wrong directory depth (non-canonical layout) |
wipe | Remove all regeneratable outputs so scripts can be re-run from scratch |
sync | Rsync the remote analyses tree to a local directory |
import-stats | Import existing CSV statistics files into statistics.db |
clean
Removes files that do not conform to the canonical layout — for example, plots
written directly under validations/<experiment>/plots/ instead of the correct
plots/<domain>/<model>/<period>/<type>/ path. Canonical files are left
untouched.
# Preview what would be removed
manage-analyses clean
# Remove non-conforming files
manage-analyses clean --apply
Run clean after migrating from an older layout version to sweep up any
residual artefacts that fix_analyses_layout moved but did not delete.
wipe
Removes all regeneratable outputs — PNG plots and statistics files — from the entire analyses tree so that all validation scripts can be re-run cleanly.
Non-regeneratable content is never touched:
.mdnarrative files (hand-written area descriptions, experiment notes)metadata.yamlsimulation_list.dbandstaging/
# Preview what would be removed
manage-analyses wipe
# Wipe all regeneratable outputs
manage-analyses wipe --apply
Regeneratable patterns removed by wipe:
*_validation_statistics.txt
*_3d_validation_statistics.txt
*_profile_validation_statistics.txt
*_tidal_statistics.txt
*_tidal_detailed_statistics.txt
*_scenario_statistics.txt
gesla_station_comparison.csv
station_comparison.csv
*_validation_statistics.csv
*_trends.txt
**/*.png
sync
Rsyncs the remote analyses tree to a local directory. Only files that conform to the canonical layout are transferred; old-layout artefacts at the wrong directory depth are automatically excluded via rsync filter rules.
# Dry-run sync (shows what would transfer — safe to run first)
manage-analyses sync kb@remote:/data/analyses/
# Actual sync
manage-analyses sync kb@remote:/data/analyses/ --apply
# Sync to a specific local directory
manage-analyses sync kb@remote:/data/analyses/ \
--analyses-dir /data/local/analyses --apply
# Print the rsync command only (for manual inspection or tweaking)
manage-analyses sync kb@remote:/data/analyses/ --print-cmd
# Use a specific SSH key
manage-analyses sync kb@remote:/data/analyses/ --apply --ssh ~/.ssh/id_rsa
update-from-remote calls manage-analyses sync internally. Use
manage-analyses sync directly when you want to sync without triggering the
full report/deploy pipeline.
import-stats
Walks the analyses/ tree for all *_validation_statistics.csv files and
upserts their rows into analyses/statistics.db. This is a one-off migration
for results produced before the database was introduced; new runs populate the
database automatically.
# Preview what would be imported (dry run)
manage-analyses import-stats
# Actually import
manage-analyses import-stats --apply
The importer infers area, experiment, domain, and model from the file’s directory path. See the Statistics Database guide for details on the schema and how to query the database.
–analyses-dir
All subcommands accept --analyses-dir DIR to set the local analyses root.
There is no ./analyses fallback: when omitted, the root is expanded from
OCEANICU_ANALYSES_FOLDER in this machine’s data-roots file
(<hostname>_ocean-post_data_roots.yaml). If neither is set, the command
stops with an error naming the missing variable — wipe in particular is
destructive, so it never guesses a directory from the current path.
Canonical layout reminder
The canonical layout that clean and sync enforce (defined in lib/layout.py):
<analyses_dir>/areas/<AREA>/validations/<EXPERIMENT>/
plots/<domain>/<model>/<period>/<type>/
tables/<domain>/<model>/<type>/
Where:
<domain>isphysicsorbio<model>is the model name, e.g.pyGETM<period>is a year (2020) or a range (2015-2022)<type>issurface,bottom,3d,argo,tidal,wod,cruise,platform,ices, or a depth slice like0050m
<EXPERIMENT> itself can be a single name (the legacy layout, still used by
NS and AMM7) or <SOURCE>/<experiment> — e.g. CMEMS/tidal,
CMIP6_raw/GFDL-ESM4-ssp126/run01 (NSe’s layout, selected via --source on
the validation scripts). Both forms are recognised by discovery and by
clean/sync; see sources.yaml
and the tidal/gridded guides for how the --source form is produced.