合并Pentad NetCDF文件时时间维度日期覆盖问题求助
合并Pentad NetCDF文件时时间维度覆盖问题的解决方法
问题描述
我有1482个命名格式为precip_yyyymmd.nc的Pentad NetCDF文件(示例:precip_2000061.nc),其中precip_为通用前缀,yyyy代表年份,mm代表月份,d代表当月的Pentad编号(每月共6个),时间跨度为2000年6月至2020年12月。合并后期望time维度包含1482个日期(格式如2000-06-01、2000-06-02……2020-12-06),但现有脚本运行后日期出现相互覆盖,time维度未生成全部1482个日期。
原代码
import xarray as xr import os # set the directory path where the files are stored directory = '/home/wilson/Documents/PYTHON_TUTORIAL/Aggregated_Data_sum' # get a list of all netCDF files in the directory file_list = [f for f in os.listdir(directory) if f.endswith('.nc')] # create an empty list to store the xarray datasets datasets = [] # loop over each file in the file list for file in file_list: # load the netCDF file into an xarray dataset ds = xr.open_dataset(os.path.join(directory, file)) # extract the year, month, and pentad from the file name year=int(file[7:11]) month=int(file[11:13]) pentad=int(file[13:14]) # add the year, month, and pentad as a combined time coordinate ds = ds.expand_dims(time=[f"{year:04d}-{month:02d}-{pentad:02d}"]) # append the dataset to the list datasets.append(ds) # concatenate the list of datasets into a single xarray dataset along the 'time' dim merged_ds = xr.concat(datasets, dim='time') # sort the merged dataset by the 'time' dimension merged_ds = merged_ds.sortby('time') # save the merged dataset as a new netCDF file merged_ds.to_netcdf(os.path.join(directory, f"merged_pentad.nc"))
单文件输出信息
<xarray.Dataset> Dimensions: (Lon: 40, Lat: 20, time: 1) Coordinates: * time (time) <U10 '2019-10-06' * Lon (Lon) float64 -19.5 -18.5 -17.5 -16.5 -15.5 ... 16.5 17.5 18.5 19.5 * Lat (Lat) float64 0.5 1.5 2.5 3.5 4.5 5.5 ... 15.5 16.5 17.5 18.5 19.5 Data variables: precip (time, Lat, Lon) float32 nan nan nan nan nan ... 0.0 0.0 0.0 0.0
问题原因
- 文件名提取逻辑存在局限性:原代码中
pentad=int(file[13:14])仅提取单个字符,若Pentad编号为多位数(虽当前场景为1-6,但逻辑不通用)会导致提取错误,进而生成重复时间坐标。 - 时间坐标类型错误:使用字符串作为时间坐标,xarray无法正确识别为时间序列,合并或排序时易出现覆盖、排序异常等问题。
- 文件列表无时间顺序:
os.listdir()返回的文件顺序不固定,若未提前排序,可能导致后续合并后时间维度混乱。
修复后的代码
import xarray as xr import os import pandas as pd # 设置文件存储目录 directory = '/home/wilson/Documents/PYTHON_TUTORIAL/Aggregated_Data_sum' # 获取所有NetCDF文件 file_list = [f for f in os.listdir(directory) if f.endswith('.nc')] # 自定义排序规则:按文件名中的年、月、Pentad编号排序 def sort_by_time(filename): year = int(filename[7:11]) month = int(filename[11:13]) pentad = int(filename[13:-3]) # 从第13位取到.nc前,兼容多位数Pentad return (year, month, pentad) # 对文件列表按时间排序 file_list.sort(key=sort_by_time) datasets = [] for file in file_list: ds = xr.open_dataset(os.path.join(directory, file)) # 提取时间信息 year = int(file[7:11]) month = int(file[11:13]) pentad = int(file[13:-3]) # 生成datetime类型的时间坐标 time_dt = pd.to_datetime(f"{year:04d}-{month:02d}-{pentad:02d}") # 添加time维度并设置为datetime类型 ds = ds.expand_dims(time=[time_dt]) datasets.append(ds) # 沿time维度合并数据集 merged_ds = xr.concat(datasets, dim='time') # 保存合并后的文件 merged_ds.to_netcdf(os.path.join(directory, "merged_pentad.nc")) # 验证结果 print(f"合并后的time维度长度:{len(merged_ds.time)}") print(f"时间范围:{merged_ds.time.min().values} 至 {merged_ds.time.max().values}")
关键修复说明
- 优化文件名提取:使用
file[13:-3]提取Pentad编号,兼容任意长度的编号,避免索引错误导致的时间重复。 - 转换时间坐标类型:通过
pd.to_datetime()将字符串时间转换为datetime64类型,xarray可正确识别时间序列,避免合并时的覆盖问题。 - 提前排序文件列表:按时间顺序处理文件,确保合并后的数据集时间维度顺序正确,减少后续排序开销。
内容的提问来源于stack exchange,提问作者hilorywilsmart
相关产品推荐
相关产品推荐

