update-from-remote Guide

Overview

update-from-remote is a one-command post-simulation pipeline. After a model run finishes on a remote machine and the analysis scripts have written their outputs there, this command:

  1. sync — rsyncs the remote analyses/ tree to the local machine
  2. merge — flushes staging YAML records into simulation_list.db
  3. scan — registers any experiments found on the filesystem but not yet in the DB
  4. report — regenerates all Hugo pages via ocean-reporting
  5. deploy — builds and pushes the Hugo site to gh-pages (only with --deploy)

Both REMOTE and LOCAL_ANALYSES are required positional arguments. There are no defaults, which prevents accidentally syncing into the wrong directory.

Basic usage

# Safe first run: dry-run sync, all other steps shown but skipped
update-from-remote kb@myhost:/data/analyses/ ./analyses

# Sync + merge + report (no deploy)
update-from-remote kb@myhost:/data/analyses/ ./analyses --apply

# Full pipeline: sync + merge + report + deploy to gh-pages
update-from-remote kb@myhost:/data/analyses/ ./analyses --apply --deploy

# Sync + merge + report + local preview
update-from-remote kb@myhost:/data/analyses/ ./analyses --apply --serve

# Skip sync (only re-run report/deploy from existing local data)
update-from-remote kb@myhost:/data/analyses/ ./analyses --no-sync --apply

# Dry-run everything: print all commands without executing any
update-from-remote kb@myhost:/data/analyses/ ./analyses --dry-run

Options

FlagDescription
REMOTERemote analyses path, e.g. kb@myhost:/data/analyses/
LOCAL_ANALYSESLocal analyses directory, e.g. ./analyses
--applyActually execute each step (default: sync shows diff, others are skipped)
--dry-runPrint all commands without executing any step
--no-syncSkip the rsync step
--no-reportSkip the ocean-reporting step
--deployAfter reporting, build and push Hugo to gh-pages
--serveAfter reporting, start a local Hugo preview server
--hugo-dir DIRPath to Hugo site directory
--ssh KEYSSH private key for rsync (-i flag)

–apply vs –dry-run

Without --apply, only the rsync step runs (in dry-run mode) so you can see what would transfer. All other steps (merge, scan, report, deploy) are printed but skipped.

With --apply, all steps execute. The deploy or serve step only runs if reporting succeeds.

--dry-run overrides everything: no step executes at all, only the commands are printed.

Sync behaviour

The sync uses manage-analyses sync internally, which only transfers files that conform to the canonical analyses/ layout. Old-layout artefacts at incorrect directory depths are excluded automatically. See guides/manage-analyses.md for details.

Deploy

The deploy step requires deploy_ghpages.py at the repository root. It builds the Hugo site and force-pushes the gh-pages branch. Only runs if:

  • --deploy is passed
  • --apply is passed
  • the reporting step succeeded

Serve

--serve starts hugo server on a free local port after reporting. Useful for reviewing the updated site before deploying. Cannot be combined with --deploy.

Typical workflow

# 1. Model run finishes on the cluster

# 2. Check what would sync (safe, no data moved)
update-from-remote kb@cluster:/scratch/analyses/ ./analyses

# 3. Sync + report locally
update-from-remote kb@cluster:/scratch/analyses/ ./analyses --apply

# 4. Review at http://localhost:1313
update-from-remote kb@cluster:/scratch/analyses/ ./analyses \
    --no-sync --apply --serve

# 5. Deploy when satisfied
update-from-remote kb@cluster:/scratch/analyses/ ./analyses \
    --no-sync --apply --deploy