Regridding User Guide
ocean_data.regridding provides xESMF-based regridding with automatic
weight-file caching, domain subsetting, and coordinate name normalisation.
The regrid command (cli/regrid.py) exposes the same functionality as a
command-line tool. ocean-post re-exports the same API from
ocean_post.regridding for backward compatibility; import from
ocean_data.regridding in new code.
Contents
- Introduction
- Installation and setup
- Command-line interface
- Python API
- Weight caching
- Regridding methods
- Coordinate name handling
- Domain subsetting
- Unit conversion
- Troubleshooting
1. Introduction
Regridding maps data from one spatial grid to another — for example interpolating a global 0.25° satellite observation product onto a regional 1 km model grid, or the other way round, before computing validation statistics.
You need regridding when:
- Source and target files have different spatial resolutions or projections.
- You want to compute point-by-point differences (RMSE, bias) between two datasets that do not share a common grid.
- You are comparing a regional model against a global observation product and only want the overlapping domain.
This module provides:
RegridManager— full-featured class with weight caching, domain subsetting, and automatic coordinate name handling.regrid_data()— convenience one-liner for single-use regridding.open_with_chunks()— helper to open a NetCDF file with Dask chunking before passing it to the regridder.regrid— CLI wrapper that reads source and target files, regrids all (or selected) variables, and writes a NetCDF output file.
2. Installation and setup
Core dependencies (always required):
The ocean-stack conda env (see ../ocean-stack/environment.yml) already
installs xESMF/ESMF from conda-forge and all four ocean-* packages in
editable mode:
conda activate ocean-stack
xESMF depends on ESMF, which is most reliably installed via conda — if
building an env manually rather than from ocean-stack:
conda install -c conda-forge esmf esmpy xesmf
pip install xarray numpy
Optional — Dask (for chunked/out-of-core processing):
pip install dask
If xesmf is not installed, importing RegridManager raises an ImportError
with install instructions. The module can still be imported; only calls that
construct a RegridManager will fail.
3. Command-line interface
Run regrid --help for a full flag listing. All INFO
messages go to stderr; only the xarray repr is printed to stdout.
3.1 Basic usage
Three arguments are always required: --source, --target, and --output.
regrid \
--source /data/obs/sst.nc \
--target /data/model/output.nc \
--output regridded_sst.nc
The source file is opened via DataLoader (supporting the same fallback
time-decode chain as ocean_data.data_loader). All data variables in the source
file are regridded by default. The target file is opened with plain
xr.open_dataset; only its coordinates are used — target values are ignored.
The output file is a NetCDF dataset with three extra global attributes recording provenance:
regrid_method bilinear
regrid_source_file /data/obs/sst.nc
regrid_target_file /data/model/output.nc
3.2 Selecting variables
By default every data variable in the source file is regridded. Use
--variable to restrict to one or more variables:
# Single variable
regrid \
--source /data/obs/sst.nc \
--target /data/model/output.nc \
--output regridded_sst.nc \
--variable analysed_sst
# Multiple variables, comma-separated (no spaces around commas)
regrid \
--source /data/obs/ts_profiles.nc \
--target /data/model/grid.nc \
--output regridded_ts.nc \
--variable TEMP,PSAL
If a requested variable is not present in the source file the script exits with an error that lists the available names.
3.3 Choosing the regrid method
Use --method to select the interpolation algorithm. The default is
bilinear.
regrid \
--source /data/obs/sst.nc \
--target /data/model/output.nc \
--output regridded_sst.nc \
--method conservative
See Regridding methods for a full comparison.
3.4 The –target-variable flag
The CLI reads coordinate arrays from a variable in the target file. By
default it uses the first data variable it finds. Use --target-variable to
specify a different one — useful when the first variable has an unusual grid
or when the target file contains multiple grids:
regrid \
--source /data/obs/sst.nc \
--target /data/model/multi_grid.nc \
--output regridded_sst.nc \
--target-variable temperature
If the specified variable is not found the script exits with an error listing available names.
3.5 Weight caching
Weight files are stored in ./regrid_weights/ by default and reused
automatically on subsequent calls with the same source/target grids and
method. Use --cache-dir to put them elsewhere, or --no-cache to disable
caching entirely:
# Custom cache directory
regrid \
--source /data/obs/sst.nc \
--target /data/model/output.nc \
--output regridded_sst.nc \
--cache-dir /scratch/weights
# Disable caching (always recompute weights)
regrid \
--source /data/obs/sst.nc \
--target /data/model/output.nc \
--output regridded_sst.nc \
--no-cache
See Weight caching for details on how the cache key is computed.
3.6 Global source grids
When the source grid wraps around the globe (e.g. OISST, ERA5), pass
--periodic so xESMF treats the longitude boundary as periodic rather than
as an open edge:
regrid \
--source /data/global/oisst.nc \
--target /data/model/NS.nc \
--output oisst_on_NS_grid.nc \
--periodic
Without --periodic, points near the dateline or prime meridian can be
assigned NaN because xESMF does not look across the boundary.
3.7 Skipping domain subsetting
Before regridding, the source is automatically subset to the target domain
plus a 2° buffer (see Domain subsetting). Pass
--no-subset to disable this step:
regrid \
--source /data/obs/sst.nc \
--target /data/model/output.nc \
--output regridded_sst.nc \
--no-subset
Use --no-subset if the source and target grids cover the same region (no
memory saving is possible), or when subsetting produces edge artefacts with
the conservative method near the domain boundary.
4. Python API
ocean_data is an installed package (pip install -e ../ocean-data, see
../ocean-stack/environment.yml), so a plain import is enough:
from ocean_data.regridding import RegridManager, regrid_data, open_with_chunks
4.1 RegridManager
The main class. Instantiate once per session; the object holds the cache
directory path and a CoordinateMapper instance that is reused across calls.
from ocean_data.regridding import RegridManager
manager = RegridManager(
cache_dir='./regrid_weights', # directory for weight files
verbose=True, # print progress messages
)
Constructor parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
cache_dir | str | ~/.cache/ocean-data/regridding | Directory where weight NetCDF files are stored. Created automatically if absent. The regrid CLI overrides this default to ./regrid_weights via its own --cache-dir flag. |
verbose | bool | True | Enable INFO-level progress output. |
RegridManager raises ImportError at construction time if xESMF is not
installed.
4.2 .regrid()
Regrids a source DataArray onto the grid of a target DataArray.
import xarray as xr
source_ds = xr.open_dataset('/data/obs/sst.nc')
target_ds = xr.open_dataset('/data/model/output.nc')
result = manager.regrid(
source_ds['analysed_sst'],
target_ds['temperature'],
)
Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
source | xr.DataArray | — | Data to regrid. |
target | xr.DataArray | — | Provides the output grid; its values are ignored. |
method | str | 'bilinear' | Interpolation method. See Regridding methods. |
subset | bool | True | Subset source to target domain before regridding. |
periodic | bool | False | Treat longitude as periodic (use for global grids). |
reuse_weights | bool | True | Load cached weights if they exist; save new ones when they do not. |
.regrid() performs no unit conversion — see
Unit conversion.
Return value: an xr.DataArray on the target grid with the same name and
attributes as the source.
# Bilinear interpolation with weight caching
sst_on_model_grid = manager.regrid(
source_ds['analysed_sst'],
target_ds['temperature'],
method='bilinear',
)
# Conservative — better for fluxes; disable subsetting to avoid edge effects
flux_on_model_grid = manager.regrid(
source_ds['net_heat_flux'],
target_ds['temperature'],
method='conservative',
subset=False,
)
# Nearest-neighbour, no caching
mask_on_model_grid = manager.regrid(
source_ds['sea_ice_mask'],
target_ds['temperature'],
method='nearest_s2d',
reuse_weights=False,
)
Note:
methodmust be one of the five supported values listed in Regridding methods. Passing an unknown name raises aValueError.
4.3 .subset_to_target_domain()
Subsets a source DataArray to the bounding box of a target DataArray plus
an optional buffer in degrees. Called automatically by .regrid() when
subset=True, but also available independently.
source_subset = manager.subset_to_target_domain(
source_ds['sst'],
target_ds['temperature'],
buffer=2.0, # degrees of padding around target domain
)
Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
source | xr.DataArray | — | Data to subset. |
target | xr.DataArray | — | Defines the bounding box. |
buffer | float | 2.0 | Extra degrees of padding on each side. |
The method resolves the lat/lon coordinate names via CoordinateMapper
(falling back to a hardcoded alias list if name_mappings.yaml isn’t
available), then uses .sel() with a slice on the lat/lon dimension
coordinates — sorting first if a dimension is in descending order, since
.sel(slice(...)) silently returns zero results otherwise. If no
recognized lat/lon dimension is found, the source is returned unchanged and
a debug message is logged; if .sel() itself fails for any other reason a
warning is logged and the source is also returned unchanged.
4.4 regrid_data()
A convenience wrapper that creates a throwaway RegridManager and calls
.regrid() in a single step. Suitable for one-off use in scripts where you
do not need to reuse the manager across multiple calls.
from ocean_data.regridding import regrid_data
result = regrid_data(
source_ds['sst'],
target_ds['temperature'],
method='bilinear',
subset=True,
cache_dir='./regrid_weights',
verbose=False,
)
Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
source | xr.DataArray | — | Data to regrid. |
target | xr.DataArray | — | Target grid reference. |
method | str | 'bilinear' | Interpolation method. |
subset | bool | True | Subset source to target domain. |
cache_dir | str | ~/.cache/ocean-data/regridding | Weight file directory. |
verbose | bool | False | Progress output. |
Return value: an xr.DataArray on the target grid.
Note:
regrid_data()creates a newRegridManageron every call. For repeated regridding operations (e.g. in a loop over time periods), use a singleRegridManagerinstance directly so theCoordinateMapperand directory handle are reused.
4.5 open_with_chunks()
Opens a NetCDF file and returns a single variable as a Dask-backed
DataArray with time chunking. Useful when the source file is too large to
fit in memory.
from ocean_data.regridding import open_with_chunks
sst = open_with_chunks(
'/data/obs/sst_2023.nc',
variable='analysed_sst',
time_chunk=30, # 30 time steps per Dask chunk
)
result = manager.regrid(sst, target_ds['temperature'])
Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
path | str | — | Path to the NetCDF file. |
variable | str | — | Variable name to extract. |
time_chunk | int | 30 | Number of time steps per Dask chunk. |
Return value: xr.DataArray backed by Dask (lazy, not yet loaded).
5. Weight caching
Computing xESMF interpolation weights is the most expensive part of regridding. For a pair of grids the weights need to be computed only once; subsequent calls with the same grids reuse the cached file, reducing wall time from minutes to seconds.
How the cache key is computed:
The weight filename is derived from the first 12 hex characters of an MD5 hash of the following string:
{source.shape}_{target.shape}_{method}_{src_lat_min}_{src_lat_max}_{src_lon_min}_{src_lon_max}_{tgt_lat_min}_{tgt_lat_max}_{tgt_lon_min}_{tgt_lon_max}
Array shape plus the full lat/lon bounding box (min and max, both axes) of each grid are included in the key — not just the latitude minimum — so two grids with the same number of points but covering different geographic regions never share weights. If the lat/lon bounds can’t be extracted for some reason, the key silently falls back to shape + method only.
Weight files are saved as NetCDF (xESMF native format) under the cache directory:
regrid_weights/
└── weights_bilinear_a3f9c12de801.nc
└── weights_conservative_1b4e7f2a9c03.nc
Cache behaviour:
- On first call: weights are computed and saved.
- On subsequent calls with the same grids: the existing file is loaded.
reuse_weights=False(or--no-cache): weights are neither loaded nor saved.
Note: If you change the spatial extent of your data between runs but keep the same array shape, the hash will change and a new weight file will be created — the old one is not deleted automatically. Clean up
./regrid_weights/periodically if you are experimenting with different domains.
6. Regridding methods
| Method | Flag value | Description | Best for |
|---|---|---|---|
| Bilinear | bilinear | Distance-weighted average of surrounding source points. Fast and smooth. | General purpose; SST, SSH, temperature fields |
| Conservative | conservative | Preserves the area-weighted integral. Slower than bilinear. | Fluxes, precipitation, quantities that must be budgeted |
| Nearest source to dest | nearest_s2d | Each destination point takes the value of its nearest source point. | Categorical masks, flag variables, binary fields |
| Nearest dest to source | nearest_d2s | Each source point is assigned to its nearest destination point. | Upsampling (coarse → fine), filling in sparse observations |
| Patch | patch | Higher-order conservative. Smoother than conservative but slower. | High-accuracy conservative regridding where smoothness matters |
The default is bilinear. For most ocean validation use cases — SST, SSS,
SSH — bilinear is the right choice. Use conservative when you need to
preserve area integrals, for example when regridding heat fluxes before
computing a budget.
7. Coordinate name handling
xESMF requires that the latitude and longitude coordinate arrays in both
source and target DataArrays be named exactly lat and lon. Many ocean
model formats use different names:
| Model | Latitude coord | Longitude coord |
|---|---|---|
| NEMO | nav_lat | nav_lon |
| Custom (this project) | latt | lont |
| ROMS | lat_rho | lon_rho |
| CESM | TLAT | TLONG |
| ERA5 | latitude | longitude |
RegridManager._prepare_for_xesmf() handles this automatically. Before
building the xESMF regridder it uses CoordinateMapper (from
ocean_data.name_mappings, driven by ocean-data/config/name_mappings.yaml)
to find the actual geographic coordinate names, and handles the two cases
differently:
Dimension coordinate (e.g. ERA5’s
latitude/longitude, which are also dimension names): left alone. xESMF’s CF accessor already finds it bystandard_name— adding alatalias on top would create a second coordinate with the samestandard_nameand make xESMF raise “multiple variables for key ’latitude’” during conservative regridding.Non-dimension coordinate (e.g. NEMO’s
nav_lat/nav_lon, this project’slatt/lont, both 2-D arrays indexed byy/x): alat/lonalias is assigned and the original coordinate name is dropped:# Internally, for a NEMO file with a 2-D nav_lat coordinate: data = data.assign_coords(lat=data.coords['nav_lat']).drop_vars('nav_lat')
_prepare_for_xesmf() also sorts any descending 1-D lat/lon dimension
axis to ascending (xESMF requires monotonically increasing coordinates) and
converts a non-contiguous array to C-contiguous (xESMF otherwise warns and
runs slower).
This happens transparently — you never need to rename coordinates yourself.
To add support for a new model format, add its coordinate names to
ocean-data/config/name_mappings.yaml under the latitude and longitude
alias lists. No Python changes are needed.
Note: for a non-dimension coordinate, the original name (
nav_lat,latt, etc.) is dropped, not kept alongside the newlat/lonalias — the regridded result only haslat/lon. For a dimension coordinate like ERA5’slatitude/longitude, nothing is renamed or dropped at all; the original name is what xESMF actually uses.
8. Domain subsetting
When the source grid covers a much larger area than the target (for example regridding a global observation product onto a small regional model), loading the full source into memory before regridding is wasteful. The default behaviour is to subset the source first:
# The regrid() call automatically subsets when subset=True (the default)
result = manager.regrid(source_da, target_da, subset=True)
# Which is equivalent to calling these two steps manually:
source_sub = manager.subset_to_target_domain(source_da, target_da, buffer=2.0)
result = manager.regrid(source_sub, target_da, subset=False)
The buffer parameter (default 2.0°) adds a margin around the target
bounding box before subsetting. This ensures that interpolation stencils
near the domain edges have enough surrounding source points to produce
accurate results.
When to use --no-subset / subset=False:
- Source and target cover the same region (no data would be discarded).
- The source grid is small and fits in memory anyway.
- Conservative regridding near the domain edge produces artefacts because the clipped source boundary introduces sharp gradients — in this case load the full source and let xESMF use its natural boundary conditions.
Note: Subsetting uses
.sel()with a slice on the lat/lon dimension (sorted to ascending order first if needed), using bounds read from the actual geographic coordinate array (resolved viaCoordinateMapper, so it works whether that array islat,nav_lat,latt, etc.). If the dimension can’t be identified at all, or.sel()fails for any reason (e.g. the dimension has no attached index coordinate to select on, which can happen for some 2-D-coordinate grids), subsetting is skipped: a warning is logged and the source is returned unchanged rather than raising. For such files usesubset=Falseexplicitly to skip the attempt.
9. Unit conversion
RegridManager.regrid() performs no unit conversion of its own — it
regrids values exactly as given, in whatever units the source happens to be
in. If the source and target are in different units (e.g. temperature in
Kelvin vs Celsius) the regridded output stays in the source’s original
units; convert before or after calling .regrid().
One existing convention for this, used by gridded-2d-validation’s own
observation-loading config (not by regrid/RegridManager itself), is a
per-dataset offset/scale_factor pair next to the dataset path:
observations:
sst:
CCI-SST:
path: "${CCI_SST_FOLDER}/sst_{year}.nc"
offset: -273.15 # Kelvin → Celsius, applied after loading
That mechanism is specific to the observation-loading step of the
validation pipeline, not a general feature of ocean_data.regridding — if
you’re calling RegridManager/regrid_data directly, apply any unit
conversion to the source DataArray yourself before passing it in.
10. Troubleshooting
ImportError: xESMF is required
ImportError: xESMF is required. Install with: pip install xesmf
Install xESMF via conda for the most reliable ESMF build:
conda install -c conda-forge esmf esmpy xesmf
ValueError: Method must be one of [...]
The --method value (or method argument) is not one of the five supported
names. Check the Regridding methods table.
ValueError: --target-variable 'X' not found in target file
The variable named by --target-variable does not exist in the target file.
Run a quick inspection to see what variables are available:
data-loader /data/model/output.nc --info
Output is all NaN
This usually means xESMF could not find geographic coordinates. The most
common cause is that _prepare_for_xesmf() could not map the coordinate
names because they are not in ocean-data/config/name_mappings.yaml.
Check what coordinate names the file uses:
data-loader /data/model/output.nc --info
Then add the missing names to ocean-data/config/name_mappings.yaml under
latitude or longitude.
Output has data only in part of the domain / edge NaN strip
The source does not fully cover the target domain after subsetting, or the buffer was too small. Try:
- Increasing the buffer by calling
.subset_to_target_domain()manually with a largerbuffervalue before passing to.regrid(subset=False). - Passing
--no-subset/subset=Falseto use the full source extent. - For global sources that wrap around the dateline, add
--periodic/periodic=True.
Conservative regridding raises an ESMF error about grid bounds
The conservative method requires cell-corner coordinates (bounds), not just
cell-centre coordinates. If the source or target file does not contain bounds
(e.g. lat_bnds, lon_bnds), xESMF falls back to estimating them, which
may fail for irregular grids. Switch to bilinear or nearest_s2d for
irregular-grid files that lack explicit bounds.
Weight file is stale after changing the source domain
If you re-run with a different spatial subset of the same source file but the
array shape happens to be the same as before, the hash changes (because the
coordinate extremes change) and a fresh weight file is created automatically.
If the shape also changed, the hash changes too. Stale weight files are
never automatically deleted; remove the ./regrid_weights/ directory to
start fresh.
Could not subset source to target domain warning
The source file has 2-D dimension coordinates (e.g. curvilinear NEMO grids
where y/x are the dimensions and nav_lat/nav_lon are 2-D non-dimension
coords). .sel() with a lat/lon slice cannot operate on such grids. Pass
subset=False (or --no-subset) to skip subsetting and proceed with
regridding on the full source array.