Profile Validation Guide

Overview

Five scripts share the same profile validation workflow:

CommandObservation source
argo-profilesArgo floats (Ifremer ERDDAP)
glodap-profilesGLODAP v2.2023 bottle data
wod-profilesWorld Ocean Database (NOAA ERDDAP)
cruise-ctd-profilesICES / EMODnet cruise CTD casts
ices-profilesICES/ECOVAL profiles (local Feather files)
fixed-platformEMODnet fixed-platform (mooring/buoy) profiles

All six follow the same config structure, output layout, and command-line flags.

Config structure

Annotated example configs: config/argo_profiles.yaml · config/glodap_profiles.yaml · config/cruise_ctd_profiles.yaml · config/wod_profiles.yaml · config/ices_profiles.yaml · config/fixed_platform.yaml

# config/argo_profiles.yaml (per-run)
area: NA
experiment: Baseline   # overridden by --experiment when --source is used

years: [2020, 2022]           # or year: 2021 for a single year
domain:
  lon_min: -20.0
  lon_max:  10.0
  lat_min:  40.0
  lat_max:  65.0

variables: [TEMP, PSAL, DOXY] # subset of available variables

argo:
  dataset: realtime
  wmo_numbers: []         # leave empty to download all floats in domain
  cache_dir: "${CACHE_ROOT}/argo"

model:
  # Location, layout and variable names come from config/sources.yaml
  # (use --source). See the section below. Remove the model block entirely
  # for obs-only diagnostics.
  name: "pyGETM"
  filename_pattern: "nse_3d.nc"
  layout: explicit_file
  subdir:
    CMEMS: "{experiment}"
    WOA: "{experiment}"
    CMIP6: "{model}/{scenario}/{experiment}"
    CMIP6_raw: "{model}/{scenario}/{experiment}"
  run_model: GFDL-ESM4     # CMIP6 / CMIP6_raw run: overridden by --model
  run_scenario: ssp126     # CMIP6 / CMIP6_raw run: overridden by --scenario
  variable_map:
    TEMP: temp
    PSAL: salt
    DOXY: doxy

output:
  analyses_dir: "${OCEANICU_ANALYSES_FOLDER}"   # no default — see below
  save_residuals: true     # default on; set false to disable
  prefix: "na_argo"         # stem for overview figure filenames

Required keys: area, experiment, and (only without --source) model.base_path, model.filename_pattern.

Model source: --source / --experiment / --model / --scenario

Model input comes from config/sources.yaml, selected with --source — the same mechanism as tidal and gridded validation:

argo-profiles --config config/argo_profiles.yaml --source CMEMS --experiment run01 --years 2020 2022

# CMIP6_raw run — forcing WITHOUT bias correction (e.g. ssp126)
argo-profiles --config config/argo_profiles.yaml --source CMIP6_raw --model GFDL-ESM4 --scenario ssp126 --experiment run01 --years 2020 2022

CMIP6 is bias-corrected forcing; CMIP6_raw is not — use CMIP6_raw for raw-forcing runs. --experiment is the run folder name, and the output label becomes <SOURCE>/<experiment>, or <SOURCE>/<model>-<scenario>/<experiment> for CMIP6/CMIP6_raw.

output.analyses_dir has no default

There is no ./analyses fallback. The value is expanded from OCEANICU_ANALYSES_FOLDER in this machine’s data-roots file (<hostname>_ocean-post_data_roots.yaml), or set directly with --analyses-dir. If neither is set, the run stops with an error naming the missing variable.

Common command-line flags

All six scripts accept:

FlagDescription
--config FILEYAML config file
--area NAMEOverride area
--experiment NAMEOverride experiment (the run folder when --source is used)
--source NAMEModel source from config/sources.yaml (e.g. CMEMS, WOA, CMIP6, CMIP6_raw)
--model NAME / --scenario NAMECMIP6/CMIP6_raw model and scenario
--years START ENDOverride year range
--variables VAR …Override variables
--analyses-dir DIROverride output.analyses_dir (otherwise required via OCEANICU_ANALYSES_FOLDER)
--no-stagingSkip writing to simulation registry
--dryrunPrint configuration summary and exit without running

Script-specific flags:

  • argo-profiles, wod-profiles, cruise-ctd-profiles: no extra flags
  • glodap-profiles, ices-profiles: --no-model (skip model comparison, obs diagnostics only)
  • fixed-platform: --list-platforms (survey mode), --platform INDEX_OR_LAT,LON, --min-obs N

What each run produces

<OCEANICU_ANALYSES_FOLDER>/areas/<AREA>/validations/<SOURCE>/<experiment>/
├── plots/
│   └── physics/               ← physics | bio
│       └── pyGETM/
│           └── <period>/      ← e.g. 2020-2022
│               └── argo/      ← or glodap/, wod/, cruise/, ices/, platform/
│                   ├── <prefix>_overview.png
│                   ├── <prefix>_hovmoller_TEMP.png
│                   └── <prefix>_profiles_YYYY.png
└── tables/
    └── physics/
        └── pyGETM/
            └── argo/
                └── <prefix>_profile_validation_statistics.txt

For CMIP6/CMIP6_raw, the path is .../<SOURCE>/<model>-<scenario>/<experiment>/....

The overview figure typically has 6 panels:

  • float/station/cruise map coloured by platform or time
  • temperature vs depth scatter
  • salinity vs depth scatter
  • T-S diagram
  • mean vertical profiles (normalised)
  • profile or observation count per platform

Running multiple experiments

Run the script once per experiment/run folder:

argo-profiles --config config/argo_profiles.yaml --source CMEMS --experiment run01
argo-profiles --config config/argo_profiles.yaml --source WOA --experiment run01
argo-profiles --config config/argo_profiles.yaml --source CMIP6_raw --model GFDL-ESM4 --scenario ssp126 --experiment run01

Generating residuals for MLE comparison

save_residuals is enabled by default (true). Each run writes a parquet file alongside the standard validation outputs:

<OCEANICU_ANALYSES_FOLDER>/areas/<AREA>/validations/<SOURCE>/<experiment>/argo_residuals.parquet

The parquet has columns: time, lat, lon, depth, variable, value (obs), model_value, and optionally profile_id.

To disable:

output:
  save_residuals: false

Vertical thinning to reduce autocorrelation

Adjacent depth levels in a profile are strongly correlated, which violates the MLE independence assumption. Use vertical_thinning to thin before writing the parquet:

output:
  save_residuals: true
  vertical_thinning:
    method: min_spacing   # none (default) | min_spacing | stride
    spacing_m: 25.0       # for min_spacing: min metres between kept levels
    # stride: 3           # for stride: keep every N-th depth level
  • min_spacing — greedy algorithm: walk depths sorted ascending and keep a level only when it is ≥ spacing_m below the last kept level. Handles irregular Argo/ICES spacing naturally; recommended for most use cases.
  • stride — keep every N-th unique depth within each profile. Faster but assumes levels are roughly evenly spaced.

Thinning is applied per-cast (all variables in the same cast get the same depth levels selected), so TEMP and PSAL remain matched.

See guides/mle-comparison.md for the ranking workflow.

GLODAP-specific notes

The GLODAP v2.2023 merged master file (~300 MB) is downloaded automatically on first run and cached locally. Set glodap.cache_dir in the config to control where it lands:

glodap:
  cache_dir: "${CACHE_ROOT}/glodap"
  qc_flags: [2]   # 2 = good data; 0 includes unqualified

Pass --no-model to produce observation-only diagnostics (station map, T-S diagram, profile climatology) without loading any model output.

ICES feather-file notes

ices-profiles reads pre-exported ICES/ECOVAL quality-controlled profiles from local Apache Arrow Feather files, under observations.data_dir (e.g. ${ECOVAL_FOLDER}/point/nws/all):

<observations.data_dir>/
├── temperature/
│   └── *_temperature_<year>.feather
├── salinity/
│   └── *_salinity_<year>.feather
└── …

Unlike the ERDDAP-based scripts, no network access is required — data must be exported from the ICES Oceanographic Database (or ECOVAL) in advance.

Pass --no-model to produce observation-only diagnostics.

Fixed-platform notes

fixed-platform reads EMODnet Chemistry fixed-platform (mooring/buoy) vertical profiles via ERDDAP. Unlike scattered cruise CTD casts, a platform collects repeated profiles at the same position over time, so the output includes Hovmoller diagrams (time × depth), fixed-depth time series, and seasonal depth structure.

Because a domain can contain many platforms, the workflow is a two-step survey-then-analyse:

# Step 1 — survey: list all platforms in the domain, no heavy analysis
fixed-platform --config config/fixed_platform.yaml --list-platforms

# Step 2 — analyse: pick a platform by rank from the inventory table...
fixed-platform --config config/fixed_platform.yaml --platform 3

# ...or by coordinates
fixed-platform --config config/fixed_platform.yaml --platform 51.3,2.9

--min-obs N pre-filters platform clusters with fewer than N observations before ranking.