Regridding User Guide

ocean_data.regridding provides xESMF-based regridding with automatic weight-file caching, domain subsetting, and coordinate name normalisation. The regrid command (cli/regrid.py) exposes the same functionality as a command-line tool. ocean-post re-exports the same API from ocean_post.regridding for backward compatibility; import from ocean_data.regridding in new code.


Contents

  1. Introduction
  2. Installation and setup
  3. Command-line interface
  4. Python API
  5. Weight caching
  6. Regridding methods
  7. Coordinate name handling
  8. Domain subsetting
  9. Unit conversion
  10. Troubleshooting

1. Introduction

Regridding maps data from one spatial grid to another — for example interpolating a global 0.25° satellite observation product onto a regional 1 km model grid, or the other way round, before computing validation statistics.

You need regridding when:

  • Source and target files have different spatial resolutions or projections.
  • You want to compute point-by-point differences (RMSE, bias) between two datasets that do not share a common grid.
  • You are comparing a regional model against a global observation product and only want the overlapping domain.

This module provides:

  • RegridManager — full-featured class with weight caching, domain subsetting, and automatic coordinate name handling.
  • regrid_data() — convenience one-liner for single-use regridding.
  • open_with_chunks() — helper to open a NetCDF file with Dask chunking before passing it to the regridder.
  • regrid — CLI wrapper that reads source and target files, regrids all (or selected) variables, and writes a NetCDF output file.

2. Installation and setup

Core dependencies (always required):

The ocean-stack conda env (see ../ocean-stack/environment.yml) already installs xESMF/ESMF from conda-forge and all four ocean-* packages in editable mode:

conda activate ocean-stack

xESMF depends on ESMF, which is most reliably installed via conda — if building an env manually rather than from ocean-stack:

conda install -c conda-forge esmf esmpy xesmf
pip install xarray numpy

Optional — Dask (for chunked/out-of-core processing):

pip install dask

If xesmf is not installed, importing RegridManager raises an ImportError with install instructions. The module can still be imported; only calls that construct a RegridManager will fail.


3. Command-line interface

Run regrid --help for a full flag listing. All INFO messages go to stderr; only the xarray repr is printed to stdout.

3.1 Basic usage

Three arguments are always required: --source, --target, and --output.

regrid \
    --source /data/obs/sst.nc \
    --target /data/model/output.nc \
    --output regridded_sst.nc

The source file is opened via DataLoader (supporting the same fallback time-decode chain as ocean_data.data_loader). All data variables in the source file are regridded by default. The target file is opened with plain xr.open_dataset; only its coordinates are used — target values are ignored.

The output file is a NetCDF dataset with three extra global attributes recording provenance:

regrid_method      bilinear
regrid_source_file /data/obs/sst.nc
regrid_target_file /data/model/output.nc

3.2 Selecting variables

By default every data variable in the source file is regridded. Use --variable to restrict to one or more variables:

# Single variable
regrid \
    --source /data/obs/sst.nc \
    --target /data/model/output.nc \
    --output regridded_sst.nc \
    --variable analysed_sst

# Multiple variables, comma-separated (no spaces around commas)
regrid \
    --source /data/obs/ts_profiles.nc \
    --target /data/model/grid.nc \
    --output regridded_ts.nc \
    --variable TEMP,PSAL

If a requested variable is not present in the source file the script exits with an error that lists the available names.

3.3 Choosing the regrid method

Use --method to select the interpolation algorithm. The default is bilinear.

regrid \
    --source /data/obs/sst.nc \
    --target /data/model/output.nc \
    --output regridded_sst.nc \
    --method conservative

See Regridding methods for a full comparison.

3.4 The –target-variable flag

The CLI reads coordinate arrays from a variable in the target file. By default it uses the first data variable it finds. Use --target-variable to specify a different one — useful when the first variable has an unusual grid or when the target file contains multiple grids:

regrid \
    --source /data/obs/sst.nc \
    --target /data/model/multi_grid.nc \
    --output regridded_sst.nc \
    --target-variable temperature

If the specified variable is not found the script exits with an error listing available names.

3.5 Weight caching

Weight files are stored in ./regrid_weights/ by default and reused automatically on subsequent calls with the same source/target grids and method. Use --cache-dir to put them elsewhere, or --no-cache to disable caching entirely:

# Custom cache directory
regrid \
    --source /data/obs/sst.nc \
    --target /data/model/output.nc \
    --output regridded_sst.nc \
    --cache-dir /scratch/weights

# Disable caching (always recompute weights)
regrid \
    --source /data/obs/sst.nc \
    --target /data/model/output.nc \
    --output regridded_sst.nc \
    --no-cache

See Weight caching for details on how the cache key is computed.

3.6 Global source grids

When the source grid wraps around the globe (e.g. OISST, ERA5), pass --periodic so xESMF treats the longitude boundary as periodic rather than as an open edge:

regrid \
    --source /data/global/oisst.nc \
    --target /data/model/NS.nc \
    --output oisst_on_NS_grid.nc \
    --periodic

Without --periodic, points near the dateline or prime meridian can be assigned NaN because xESMF does not look across the boundary.

3.7 Skipping domain subsetting

Before regridding, the source is automatically subset to the target domain plus a 2° buffer (see Domain subsetting). Pass --no-subset to disable this step:

regrid \
    --source /data/obs/sst.nc \
    --target /data/model/output.nc \
    --output regridded_sst.nc \
    --no-subset

Use --no-subset if the source and target grids cover the same region (no memory saving is possible), or when subsetting produces edge artefacts with the conservative method near the domain boundary.


4. Python API

ocean_data is an installed package (pip install -e ../ocean-data, see ../ocean-stack/environment.yml), so a plain import is enough:

from ocean_data.regridding import RegridManager, regrid_data, open_with_chunks

4.1 RegridManager

The main class. Instantiate once per session; the object holds the cache directory path and a CoordinateMapper instance that is reused across calls.

from ocean_data.regridding import RegridManager

manager = RegridManager(
    cache_dir='./regrid_weights',   # directory for weight files
    verbose=True,                   # print progress messages
)

Constructor parameters:

ParameterTypeDefaultDescription
cache_dirstr~/.cache/ocean-data/regriddingDirectory where weight NetCDF files are stored. Created automatically if absent. The regrid CLI overrides this default to ./regrid_weights via its own --cache-dir flag.
verboseboolTrueEnable INFO-level progress output.

RegridManager raises ImportError at construction time if xESMF is not installed.

4.2 .regrid()

Regrids a source DataArray onto the grid of a target DataArray.

import xarray as xr

source_ds = xr.open_dataset('/data/obs/sst.nc')
target_ds = xr.open_dataset('/data/model/output.nc')

result = manager.regrid(
    source_ds['analysed_sst'],
    target_ds['temperature'],
)

Parameters:

ParameterTypeDefaultDescription
sourcexr.DataArray—Data to regrid.
targetxr.DataArray—Provides the output grid; its values are ignored.
methodstr'bilinear'Interpolation method. See Regridding methods.
subsetboolTrueSubset source to target domain before regridding.
periodicboolFalseTreat longitude as periodic (use for global grids).
reuse_weightsboolTrueLoad cached weights if they exist; save new ones when they do not.

.regrid() performs no unit conversion — see Unit conversion.

Return value: an xr.DataArray on the target grid with the same name and attributes as the source.

# Bilinear interpolation with weight caching
sst_on_model_grid = manager.regrid(
    source_ds['analysed_sst'],
    target_ds['temperature'],
    method='bilinear',
)

# Conservative — better for fluxes; disable subsetting to avoid edge effects
flux_on_model_grid = manager.regrid(
    source_ds['net_heat_flux'],
    target_ds['temperature'],
    method='conservative',
    subset=False,
)

# Nearest-neighbour, no caching
mask_on_model_grid = manager.regrid(
    source_ds['sea_ice_mask'],
    target_ds['temperature'],
    method='nearest_s2d',
    reuse_weights=False,
)

Note: method must be one of the five supported values listed in Regridding methods. Passing an unknown name raises a ValueError.

4.3 .subset_to_target_domain()

Subsets a source DataArray to the bounding box of a target DataArray plus an optional buffer in degrees. Called automatically by .regrid() when subset=True, but also available independently.

source_subset = manager.subset_to_target_domain(
    source_ds['sst'],
    target_ds['temperature'],
    buffer=2.0,      # degrees of padding around target domain
)

Parameters:

ParameterTypeDefaultDescription
sourcexr.DataArray—Data to subset.
targetxr.DataArray—Defines the bounding box.
bufferfloat2.0Extra degrees of padding on each side.

The method resolves the lat/lon coordinate names via CoordinateMapper (falling back to a hardcoded alias list if name_mappings.yaml isn’t available), then uses .sel() with a slice on the lat/lon dimension coordinates — sorting first if a dimension is in descending order, since .sel(slice(...)) silently returns zero results otherwise. If no recognized lat/lon dimension is found, the source is returned unchanged and a debug message is logged; if .sel() itself fails for any other reason a warning is logged and the source is also returned unchanged.

4.4 regrid_data()

A convenience wrapper that creates a throwaway RegridManager and calls .regrid() in a single step. Suitable for one-off use in scripts where you do not need to reuse the manager across multiple calls.

from ocean_data.regridding import regrid_data

result = regrid_data(
    source_ds['sst'],
    target_ds['temperature'],
    method='bilinear',
    subset=True,
    cache_dir='./regrid_weights',
    verbose=False,
)

Parameters:

ParameterTypeDefaultDescription
sourcexr.DataArray—Data to regrid.
targetxr.DataArray—Target grid reference.
methodstr'bilinear'Interpolation method.
subsetboolTrueSubset source to target domain.
cache_dirstr~/.cache/ocean-data/regriddingWeight file directory.
verboseboolFalseProgress output.

Return value: an xr.DataArray on the target grid.

Note: regrid_data() creates a new RegridManager on every call. For repeated regridding operations (e.g. in a loop over time periods), use a single RegridManager instance directly so the CoordinateMapper and directory handle are reused.

4.5 open_with_chunks()

Opens a NetCDF file and returns a single variable as a Dask-backed DataArray with time chunking. Useful when the source file is too large to fit in memory.

from ocean_data.regridding import open_with_chunks

sst = open_with_chunks(
    '/data/obs/sst_2023.nc',
    variable='analysed_sst',
    time_chunk=30,          # 30 time steps per Dask chunk
)

result = manager.regrid(sst, target_ds['temperature'])

Parameters:

ParameterTypeDefaultDescription
pathstr—Path to the NetCDF file.
variablestr—Variable name to extract.
time_chunkint30Number of time steps per Dask chunk.

Return value: xr.DataArray backed by Dask (lazy, not yet loaded).


5. Weight caching

Computing xESMF interpolation weights is the most expensive part of regridding. For a pair of grids the weights need to be computed only once; subsequent calls with the same grids reuse the cached file, reducing wall time from minutes to seconds.

How the cache key is computed:

The weight filename is derived from the first 12 hex characters of an MD5 hash of the following string:

{source.shape}_{target.shape}_{method}_{src_lat_min}_{src_lat_max}_{src_lon_min}_{src_lon_max}_{tgt_lat_min}_{tgt_lat_max}_{tgt_lon_min}_{tgt_lon_max}

Array shape plus the full lat/lon bounding box (min and max, both axes) of each grid are included in the key — not just the latitude minimum — so two grids with the same number of points but covering different geographic regions never share weights. If the lat/lon bounds can’t be extracted for some reason, the key silently falls back to shape + method only.

Weight files are saved as NetCDF (xESMF native format) under the cache directory:

regrid_weights/
└── weights_bilinear_a3f9c12de801.nc
└── weights_conservative_1b4e7f2a9c03.nc

Cache behaviour:

  • On first call: weights are computed and saved.
  • On subsequent calls with the same grids: the existing file is loaded.
  • reuse_weights=False (or --no-cache): weights are neither loaded nor saved.

Note: If you change the spatial extent of your data between runs but keep the same array shape, the hash will change and a new weight file will be created — the old one is not deleted automatically. Clean up ./regrid_weights/ periodically if you are experimenting with different domains.


6. Regridding methods

MethodFlag valueDescriptionBest for
BilinearbilinearDistance-weighted average of surrounding source points. Fast and smooth.General purpose; SST, SSH, temperature fields
ConservativeconservativePreserves the area-weighted integral. Slower than bilinear.Fluxes, precipitation, quantities that must be budgeted
Nearest source to destnearest_s2dEach destination point takes the value of its nearest source point.Categorical masks, flag variables, binary fields
Nearest dest to sourcenearest_d2sEach source point is assigned to its nearest destination point.Upsampling (coarse → fine), filling in sparse observations
PatchpatchHigher-order conservative. Smoother than conservative but slower.High-accuracy conservative regridding where smoothness matters

The default is bilinear. For most ocean validation use cases — SST, SSS, SSH — bilinear is the right choice. Use conservative when you need to preserve area integrals, for example when regridding heat fluxes before computing a budget.


7. Coordinate name handling

xESMF requires that the latitude and longitude coordinate arrays in both source and target DataArrays be named exactly lat and lon. Many ocean model formats use different names:

ModelLatitude coordLongitude coord
NEMOnav_latnav_lon
Custom (this project)lattlont
ROMSlat_rholon_rho
CESMTLATTLONG
ERA5latitudelongitude

RegridManager._prepare_for_xesmf() handles this automatically. Before building the xESMF regridder it uses CoordinateMapper (from ocean_data.name_mappings, driven by ocean-data/config/name_mappings.yaml) to find the actual geographic coordinate names, and handles the two cases differently:

  • Dimension coordinate (e.g. ERA5’s latitude/longitude, which are also dimension names): left alone. xESMF’s CF accessor already finds it by standard_name — adding a lat alias on top would create a second coordinate with the same standard_name and make xESMF raise “multiple variables for key ’latitude’” during conservative regridding.

  • Non-dimension coordinate (e.g. NEMO’s nav_lat/nav_lon, this project’s latt/lont, both 2-D arrays indexed by y/x): a lat/lon alias is assigned and the original coordinate name is dropped:

    # Internally, for a NEMO file with a 2-D nav_lat coordinate:
    data = data.assign_coords(lat=data.coords['nav_lat']).drop_vars('nav_lat')
    

_prepare_for_xesmf() also sorts any descending 1-D lat/lon dimension axis to ascending (xESMF requires monotonically increasing coordinates) and converts a non-contiguous array to C-contiguous (xESMF otherwise warns and runs slower).

This happens transparently — you never need to rename coordinates yourself.

To add support for a new model format, add its coordinate names to ocean-data/config/name_mappings.yaml under the latitude and longitude alias lists. No Python changes are needed.

Note: for a non-dimension coordinate, the original name (nav_lat, latt, etc.) is dropped, not kept alongside the new lat/lon alias — the regridded result only has lat/lon. For a dimension coordinate like ERA5’s latitude/longitude, nothing is renamed or dropped at all; the original name is what xESMF actually uses.


8. Domain subsetting

When the source grid covers a much larger area than the target (for example regridding a global observation product onto a small regional model), loading the full source into memory before regridding is wasteful. The default behaviour is to subset the source first:

# The regrid() call automatically subsets when subset=True (the default)
result = manager.regrid(source_da, target_da, subset=True)

# Which is equivalent to calling these two steps manually:
source_sub = manager.subset_to_target_domain(source_da, target_da, buffer=2.0)
result = manager.regrid(source_sub, target_da, subset=False)

The buffer parameter (default 2.0°) adds a margin around the target bounding box before subsetting. This ensures that interpolation stencils near the domain edges have enough surrounding source points to produce accurate results.

When to use --no-subset / subset=False:

  • Source and target cover the same region (no data would be discarded).
  • The source grid is small and fits in memory anyway.
  • Conservative regridding near the domain edge produces artefacts because the clipped source boundary introduces sharp gradients — in this case load the full source and let xESMF use its natural boundary conditions.

Note: Subsetting uses .sel() with a slice on the lat/lon dimension (sorted to ascending order first if needed), using bounds read from the actual geographic coordinate array (resolved via CoordinateMapper, so it works whether that array is lat, nav_lat, latt, etc.). If the dimension can’t be identified at all, or .sel() fails for any reason (e.g. the dimension has no attached index coordinate to select on, which can happen for some 2-D-coordinate grids), subsetting is skipped: a warning is logged and the source is returned unchanged rather than raising. For such files use subset=False explicitly to skip the attempt.


9. Unit conversion

RegridManager.regrid() performs no unit conversion of its own — it regrids values exactly as given, in whatever units the source happens to be in. If the source and target are in different units (e.g. temperature in Kelvin vs Celsius) the regridded output stays in the source’s original units; convert before or after calling .regrid().

One existing convention for this, used by gridded-2d-validation’s own observation-loading config (not by regrid/RegridManager itself), is a per-dataset offset/scale_factor pair next to the dataset path:

observations:
  sst:
    CCI-SST:
      path:   "${CCI_SST_FOLDER}/sst_{year}.nc"
      offset: -273.15      # Kelvin → Celsius, applied after loading

That mechanism is specific to the observation-loading step of the validation pipeline, not a general feature of ocean_data.regridding — if you’re calling RegridManager/regrid_data directly, apply any unit conversion to the source DataArray yourself before passing it in.


10. Troubleshooting

ImportError: xESMF is required

ImportError: xESMF is required. Install with: pip install xesmf

Install xESMF via conda for the most reliable ESMF build:

conda install -c conda-forge esmf esmpy xesmf

ValueError: Method must be one of [...]

The --method value (or method argument) is not one of the five supported names. Check the Regridding methods table.


ValueError: --target-variable 'X' not found in target file

The variable named by --target-variable does not exist in the target file. Run a quick inspection to see what variables are available:

data-loader /data/model/output.nc --info

Output is all NaN

This usually means xESMF could not find geographic coordinates. The most common cause is that _prepare_for_xesmf() could not map the coordinate names because they are not in ocean-data/config/name_mappings.yaml.

Check what coordinate names the file uses:

data-loader /data/model/output.nc --info

Then add the missing names to ocean-data/config/name_mappings.yaml under latitude or longitude.


Output has data only in part of the domain / edge NaN strip

The source does not fully cover the target domain after subsetting, or the buffer was too small. Try:

  1. Increasing the buffer by calling .subset_to_target_domain() manually with a larger buffer value before passing to .regrid(subset=False).
  2. Passing --no-subset / subset=False to use the full source extent.
  3. For global sources that wrap around the dateline, add --periodic / periodic=True.

Conservative regridding raises an ESMF error about grid bounds

The conservative method requires cell-corner coordinates (bounds), not just cell-centre coordinates. If the source or target file does not contain bounds (e.g. lat_bnds, lon_bnds), xESMF falls back to estimating them, which may fail for irregular grids. Switch to bilinear or nearest_s2d for irregular-grid files that lack explicit bounds.


Weight file is stale after changing the source domain

If you re-run with a different spatial subset of the same source file but the array shape happens to be the same as before, the hash changes (because the coordinate extremes change) and a fresh weight file is created automatically. If the shape also changed, the hash changes too. Stale weight files are never automatically deleted; remove the ./regrid_weights/ directory to start fresh.


Could not subset source to target domain warning

The source file has 2-D dimension coordinates (e.g. curvilinear NEMO grids where y/x are the dimensions and nav_lat/nav_lon are 2-D non-dimension coords). .sel() with a lat/lon slice cannot operate on such grids. Pass subset=False (or --no-subset) to skip subsetting and proceed with regridding on the full source array.