Hugo Deployment Guide

Overview

Reporting explains ocean-reporting’s options and the analyses/ tree it reads. This guide is about what happens after that: getting the generated Hugo content actually live, and the real mistakes that have already happened doing it.

There are two entry points to the exact same underlying code — pick whichever fits the situation:

  • update-from-remote — sync a remote cluster’s analyses/ down, merge/scan into the DB, then report + deploy, all in one command. Right when there’s fresh data to pull in first.
  • Manual two-step (used directly against data that’s already local, e.g. on the reporting host itself): the calling project’s own regenerate_hugo.py --apply, then this repo’s own deploy_ghpages.py --apply. Both relay over ssh to the reporting host automatically if run from elsewhere. regenerate_hugo.py isn’t part of this repo — it lives in the site’s own repo (e.g. oceanicu_3d/regenerate_hugo.py), and its own docs/web-regeneration.md has the full detail for that side.

Either path ends up calling cli.reporting:main (as ocean-reporting in one case, python3 -m cli.reporting in the other) and this repo’s own deploy_ghpages.py directly. A bug or fix in lib/reporting.py affects both equally.

deploy_ghpages.py’s actual behavior — read before assuming

It builds Hugo into its own throwaway temp directory (an explicit --destination, which overrides whatever publishDir says in the site’s Hugo config), then separately checks out gh-pages into a second temp git worktree, copies the build in, commits, and force-pushes from there. This is deliberate (see the script’s own docstring) so a stale manual hugo build never leaks into what actually goes live.

One consequence worth knowing: the site’s own configured publishDir (e.g. <site>/public/) is never touched by this pipeline at all — it’s only written by someone manually running plain hugo from the site’s Hugo directory. Don’t use its mtime as a proxy for whether the live site is current; it can sit stale indefinitely while the real deploys keep happening through the temp-dir path above.

Generated content pages: never hand-edit the output

Every page _generate_*_page-style functions in lib/reporting.py produce (e.g. _generate_nse_boundaries_page → content/areas/nse-boundaries.md) is written from scratch, every single run, from hardcoded Python string literals plus whatever plot/table files it discovers under analyses/. There is no free-text passthrough mechanism. Hand-editing one of these .md files directly survives exactly until the next report run, then is silently gone — no warning, no diff shown, nothing.

If a page needs new prose, a new table, a new section: add it to the generator function itself, not the output file. That change belongs in this repo (lib/reporting.py), even when the content is about something that lives in a different repo entirely (e.g. boundary- condition bias-correction notes for oceanicu_3d) — this repo owns every word that ends up in the generated page.

A real bug this caused, in case the pattern recurs

_generate_scenarios_page crashed with AttributeError: 'BCDiagnosticsLayout' has no attribute 'iter_all_plot_dirs' on every report run since commit 28abad5 (2026-06-23), which moved BCDiagnosticsLayout to ocean-prep as a slimmed-down per- (scenario,model) path builder and, in the process, dropped 4 reporter- side “walk the whole scenarios/ tree” methods (iter_all_plot_dirs, iter_all_table_dirs, iter_all_river_plot_dirs, iter_all_river_table_dirs) that lib/reporting.py still called.

This went unnoticed for three months because generate_hugo() calls its page-generator functions sequentially, and _generate_scenarios_page happens to run after several other pages that write successfully — so per-page output looked fine on casual inspection while the overall script silently exited non-zero every single time. Found and fixed 2026-09-21 (commit e48c882): restored the four walkers verbatim from git show 28abad5 -- lib/layout.py as local functions in lib/reporting.py itself (the on-disk path convention was unchanged since the move — only the discovery logic had been dropped). No test caught this either; tests/test_layout.py never exercised the iter_all_* methods.

If a similar “AttributeError on some Layout-ish class” turns up after a future ocean-prep refactor, check git log -p -- lib/layout.py in this repo first — it’s the same class of drift: a class moves out, gets slimmed down for its new home’s own needs, and a reporter-side helper that depended on its old, fuller surface silently breaks.

Another real bug: slashes in the experiment label broke page slugs

NSe’s current layout keys experiments as <SOURCE>/<experiment> (CMEMS/tidal, CMIP6_raw/GFDL-ESM4-ssp126/run01 — see Reporting). The page slug used to be f"{area}-{experiment}".lower(), which kept the / from the experiment label. For CMEMS/tidal that produced the slug nse-cmems/tidal, and writing the page to content/validations/nse-cmems/tidal.md raised FileNotFoundError — nse-cmems/ isn’t a directory that exists. The scenario reporter had the same bug in a second form: it keyed scenario pages by area-experiment-scenario_name, but the caller passes the same label as both experiment and scenario_name, so the slug doubled (nse-cmip6_raw/gfdl-esm4-ssp126/run01-cmip6_raw/gfdl-esm4-ssp126/run01) and nested several directories deep.

Fixed 2026-10-06 with one shared helper, validation_slug(area, experiment) in lib/layout.py, used everywhere a page slug is built in lib/tidal_reporting.py and lib/reporting.py; it turns / and _ into -. The scenario case gets its own scenario_slug(area, experiment, scenario_name) in lib/scenario_reporting.py, which drops the scenario part entirely when it repeats the experiment label, instead of doubling it. A stale nested scenario directory left over from the bug had to be removed by hand once (rm -rf content/scenarios/simulations/nse-cmip6_raw) — the generator does not clean up a directory tree it no longer writes to under the old, buggy name.

If a page slug ever contains a / or looks doubled again, the experiment or scenario label almost certainly has an unflattened / or _ in it somewhere upstream — check whether the code path still goes through validation_slug()/scenario_slug() rather than building the slug inline.

enabled_areas in regen_hosts.yaml: a publish allowlist, not a filter

The calling project’s regen_hosts.yaml (e.g. oceanicu_3d/regen_hosts.yaml) can set enabled_areas — the areas the site actually publishes. This is passed through to cli.reporting’s --area as described in Reporting: on a --recursive run, any area not listed has its already-published content//static/ pages actively removed, not left alone. Leaving an area off the list is not “not touched yet”, it is “pruned from the live site” — though nothing is lost, since the pages are always regenerated from analyses/ on request. NSe, NS and AMM7 are all enabled as of 2026-10-06; AMM7/ENA4/ENA8/NS had previously been pruned as a side effect of an --area NSe-only regen run before this allowlist existed.