NetCDF文件合并方法咨询:AVHRR Pathfinder日数据转年度文件
Hey there! As someone who’s wrangled AVHRR Pathfinder SST datasets more times than I can count, I totally get the frustration of dealing with hundreds of tiny daily .nc files. Let’s walk through reliable, efficient solutions—whether you want to fix your NCO workflow or try a more beginner-friendly tool.
Option 1: Fixing Your NCO Workflow (Since You Already Started With It)
The ncrcat command from NCO is purpose-built for concatenating NetCDF files along the time dimension, which is exactly what you need here. If you hit issues, it’s usually due to file structure mismatches or incorrect file ordering.
Step-by-Step NCO Command
Assuming your files follow a naming pattern like sst.day.YYYYMMDD.nc and sst.night.YYYYMMDD.nc, use this command to merge all daily/nightly files for a single year:
# Merge all 1981 files into one annual file (overwrites if output exists) ncrcat -O sst.*.1981*.nc 1981_annual_sst.nc
-O: Overwrites the output file if it already exists (safe to use once you’ve tested with a small subset)ncrcatautomatically sorts files by their internaltimevariable, so even if your filenames are out of order, it’ll align the data correctly.
Troubleshooting Common NCO Issues
If you ran into errors, try these checks first:
- Verify file consistency: Make sure all files have the same dimensions (lat, lon, time) and variables. Run this on a day and night file to compare:
ncdump -h sst.day.19810101.nc > day_metadata.txt ncdump -h sst.night.19810101.nc > night_metadata.txt diff day_metadata.txt night_metadata.txt - Extract matching variables/dimensions: If there are extra variables in some files, use
ncksto standardize before merging:
Then merge the standardized files instead.# Keep only the core variables needed (adjust based on your dataset) ncks -v sst,lat,lon,time sst.night.19810101.nc standardized_night.nc
Option 2: Python xarray (Beginner-Friendly & Flexible)
If you prefer writing scripts or want more control over the merging process, xarray is perfect. It’s designed for climate data and handles messy file batches gracefully.
Sample Python Script
import xarray as xr import glob # Define the year you want to process target_year = 1981 # Get all .nc files for the year, sorted by filename (or internal time) file_paths = sorted(glob.glob(f"./sst.*.{target_year}*.nc")) # Merge all files into a single dataset # `combine='by_coords'` automatically aligns data by time/lat/lon # `parallel=True` uses multiple CPU cores to speed up processing annual_dataset = xr.open_mfdataset( file_paths, combine="by_coords", parallel=True, data_vars="minimal" # Only merge variables present in all files ) # Save the merged dataset to a new NetCDF file annual_dataset.to_netcdf(f"{target_year}_annual_sst.nc")
This script is easy to adapt—just loop over target_year from 1981 to 2014 to batch-process all years at once.
Option 3: CDO (Climate Data Operators) - Blazing Fast for Large Datasets
CDO is a command-line tool built for climate data manipulation, and it’s incredibly fast for merging large batches of files. If you haven’t installed it yet, you can grab it via Conda (conda install -c conda-forge cdo).
CDO Merge Command
# Merge all 1981 files into an annual file cdo mergetime sst.*.1981*.nc 1981_annual_sst.nc
Like ncrcat, mergetime automatically sorts files by their internal time variable, so you don’t have to worry about filename order. It’s ideal if you need to process all 34 years quickly.
Quick Tips for New NetCDF Users
- Test with a small subset first: Pick 2-3 days of files to test your command/script before processing an entire year—this saves you from wasting time on a broken workflow for 365 files.
- Validate the output: Always check that the merged file has the correct time range with:
ncdump -v time 1981_annual_sst.nc | head -20 - Batch-process all years: For NCO/CDO, use a bash loop to automate all 34 years in one go:
for year in {1981..2014}; do ncrcat -O sst.*.${year}*.nc ${year}_annual_sst.nc done
内容的提问来源于stack exchange,提问作者Mihir Sharma

