Profile Validation Guide
Overview
Five scripts share the same profile validation workflow:
| Command | Observation source |
|---|---|
argo-profiles | Argo floats (Ifremer ERDDAP) |
glodap-profiles | GLODAP v2.2023 bottle data |
wod-profiles | World Ocean Database (NOAA ERDDAP) |
cruise-ctd-profiles | ICES / EMODnet cruise CTD casts |
ices-profiles | ICES/ECOVAL profiles (local Feather files) |
fixed-platform | EMODnet fixed-platform (mooring/buoy) profiles |
All six follow the same config structure, output layout, and command-line flags.
Config structure
Annotated example configs: config/argo_profiles.yaml · config/glodap_profiles.yaml · config/cruise_ctd_profiles.yaml · config/wod_profiles.yaml · config/ices_profiles.yaml · config/fixed_platform.yaml
# config/argo_profiles.yaml (per-run)
area: NA
experiment: Baseline # overridden by --experiment when --source is used
years: [2020, 2022] # or year: 2021 for a single year
domain:
lon_min: -20.0
lon_max: 10.0
lat_min: 40.0
lat_max: 65.0
variables: [TEMP, PSAL, DOXY] # subset of available variables
argo:
dataset: realtime
wmo_numbers: [] # leave empty to download all floats in domain
cache_dir: "${CACHE_ROOT}/argo"
model:
# Location, layout and variable names come from config/sources.yaml
# (use --source). See the section below. Remove the model block entirely
# for obs-only diagnostics.
name: "pyGETM"
filename_pattern: "nse_3d.nc"
layout: explicit_file
subdir:
CMEMS: "{experiment}"
WOA: "{experiment}"
CMIP6: "{model}/{scenario}/{experiment}"
CMIP6_raw: "{model}/{scenario}/{experiment}"
run_model: GFDL-ESM4 # CMIP6 / CMIP6_raw run: overridden by --model
run_scenario: ssp126 # CMIP6 / CMIP6_raw run: overridden by --scenario
variable_map:
TEMP: temp
PSAL: salt
DOXY: doxy
output:
analyses_dir: "${OCEANICU_ANALYSES_FOLDER}" # no default — see below
save_residuals: true # default on; set false to disable
prefix: "na_argo" # stem for overview figure filenames
Required keys: area, experiment, and (only without --source)
model.base_path, model.filename_pattern.
Model source: --source / --experiment / --model / --scenario
Model input comes from config/sources.yaml,
selected with --source — the same mechanism as tidal and gridded
validation:
argo-profiles --config config/argo_profiles.yaml --source CMEMS --experiment run01 --years 2020 2022
# CMIP6_raw run — forcing WITHOUT bias correction (e.g. ssp126)
argo-profiles --config config/argo_profiles.yaml --source CMIP6_raw --model GFDL-ESM4 --scenario ssp126 --experiment run01 --years 2020 2022
CMIP6 is bias-corrected forcing; CMIP6_raw is not — use CMIP6_raw for
raw-forcing runs. --experiment is the run folder name, and the output
label becomes <SOURCE>/<experiment>, or
<SOURCE>/<model>-<scenario>/<experiment> for CMIP6/CMIP6_raw.
output.analyses_dir has no default
There is no ./analyses fallback. The value is expanded from
OCEANICU_ANALYSES_FOLDER in this machine’s data-roots file
(<hostname>_ocean-post_data_roots.yaml), or set directly with
--analyses-dir. If neither is set, the run stops with an error naming the
missing variable.
Common command-line flags
All six scripts accept:
| Flag | Description |
|---|---|
--config FILE | YAML config file |
--area NAME | Override area |
--experiment NAME | Override experiment (the run folder when --source is used) |
--source NAME | Model source from config/sources.yaml (e.g. CMEMS, WOA, CMIP6, CMIP6_raw) |
--model NAME / --scenario NAME | CMIP6/CMIP6_raw model and scenario |
--years START END | Override year range |
--variables VAR … | Override variables |
--analyses-dir DIR | Override output.analyses_dir (otherwise required via OCEANICU_ANALYSES_FOLDER) |
--no-staging | Skip writing to simulation registry |
--dryrun | Print configuration summary and exit without running |
Script-specific flags:
argo-profiles,wod-profiles,cruise-ctd-profiles: no extra flagsglodap-profiles,ices-profiles:--no-model(skip model comparison, obs diagnostics only)fixed-platform:--list-platforms(survey mode),--platform INDEX_OR_LAT,LON,--min-obs N
What each run produces
<OCEANICU_ANALYSES_FOLDER>/areas/<AREA>/validations/<SOURCE>/<experiment>/
├── plots/
│ └── physics/ ← physics | bio
│ └── pyGETM/
│ └── <period>/ ← e.g. 2020-2022
│ └── argo/ ← or glodap/, wod/, cruise/, ices/, platform/
│ ├── <prefix>_overview.png
│ ├── <prefix>_hovmoller_TEMP.png
│ └── <prefix>_profiles_YYYY.png
└── tables/
└── physics/
└── pyGETM/
└── argo/
└── <prefix>_profile_validation_statistics.txt
For CMIP6/CMIP6_raw, the path is
.../<SOURCE>/<model>-<scenario>/<experiment>/....
The overview figure typically has 6 panels:
- float/station/cruise map coloured by platform or time
- temperature vs depth scatter
- salinity vs depth scatter
- T-S diagram
- mean vertical profiles (normalised)
- profile or observation count per platform
Running multiple experiments
Run the script once per experiment/run folder:
argo-profiles --config config/argo_profiles.yaml --source CMEMS --experiment run01
argo-profiles --config config/argo_profiles.yaml --source WOA --experiment run01
argo-profiles --config config/argo_profiles.yaml --source CMIP6_raw --model GFDL-ESM4 --scenario ssp126 --experiment run01
Generating residuals for MLE comparison
save_residuals is enabled by default (true). Each run writes a parquet
file alongside the standard validation outputs:
<OCEANICU_ANALYSES_FOLDER>/areas/<AREA>/validations/<SOURCE>/<experiment>/argo_residuals.parquet
The parquet has columns: time, lat, lon, depth, variable, value
(obs), model_value, and optionally profile_id.
To disable:
output:
save_residuals: false
Vertical thinning to reduce autocorrelation
Adjacent depth levels in a profile are strongly correlated, which violates
the MLE independence assumption. Use vertical_thinning to thin before
writing the parquet:
output:
save_residuals: true
vertical_thinning:
method: min_spacing # none (default) | min_spacing | stride
spacing_m: 25.0 # for min_spacing: min metres between kept levels
# stride: 3 # for stride: keep every N-th depth level
min_spacing— greedy algorithm: walk depths sorted ascending and keep a level only when it is ≥spacing_mbelow the last kept level. Handles irregular Argo/ICES spacing naturally; recommended for most use cases.stride— keep every N-th unique depth within each profile. Faster but assumes levels are roughly evenly spaced.
Thinning is applied per-cast (all variables in the same cast get the same depth levels selected), so TEMP and PSAL remain matched.
See guides/mle-comparison.md for the ranking workflow.
GLODAP-specific notes
The GLODAP v2.2023 merged master file (~300 MB) is downloaded automatically on
first run and cached locally. Set glodap.cache_dir in the config to control
where it lands:
glodap:
cache_dir: "${CACHE_ROOT}/glodap"
qc_flags: [2] # 2 = good data; 0 includes unqualified
Pass --no-model to produce observation-only diagnostics (station map, T-S
diagram, profile climatology) without loading any model output.
ICES feather-file notes
ices-profiles reads pre-exported ICES/ECOVAL quality-controlled profiles from
local Apache Arrow Feather files, under observations.data_dir
(e.g. ${ECOVAL_FOLDER}/point/nws/all):
<observations.data_dir>/
├── temperature/
│ └── *_temperature_<year>.feather
├── salinity/
│ └── *_salinity_<year>.feather
└── …
Unlike the ERDDAP-based scripts, no network access is required — data must be exported from the ICES Oceanographic Database (or ECOVAL) in advance.
Pass --no-model to produce observation-only diagnostics.
Fixed-platform notes
fixed-platform reads EMODnet Chemistry fixed-platform (mooring/buoy)
vertical profiles via ERDDAP. Unlike scattered cruise CTD casts, a platform
collects repeated profiles at the same position over time, so the output
includes Hovmoller diagrams (time × depth), fixed-depth time series, and
seasonal depth structure.
Because a domain can contain many platforms, the workflow is a two-step survey-then-analyse:
# Step 1 — survey: list all platforms in the domain, no heavy analysis
fixed-platform --config config/fixed_platform.yaml --list-platforms
# Step 2 — analyse: pick a platform by rank from the inventory table...
fixed-platform --config config/fixed_platform.yaml --platform 3
# ...or by coordinates
fixed-platform --config config/fixed_platform.yaml --platform 51.3,2.9
--min-obs N pre-filters platform clusters with fewer than N
observations before ranking.