You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用xarray.open_mfdataset预处理分配坐标报错的解决方案问询

问题:xarray open_mfdataset 结合 preprocess 分配坐标后合并失败

根据xarray文档,open_mfdataset可通过preprocess参数在合并前处理每个数据集。我的NetCDF数据集单独打开时无关联坐标,于是尝试在open_mfdataset中搭配combine='by_coords',通过自定义assignCoordinates函数预分配坐标,但运行时报错:

ValueError: Could not find any dimension coordinates to use to order the datasets for concatenation

单个数据集打开后的信息如下:

path = 'path/to/my/file/file.nc'
ds = xr.open_dataset(path, decode_times=False)
ds

# <xarray.Dataset> Size: 1GB
# Dimensions:       (comid: 2677612, time_mn: 120, time_yr: 10)
# Dimensions without coordinates: comid, time_mn, time_yr
# Data variables:
#    COMID         (comid) int32 11MB ...
#    Time_mn       (time_mn) int32 480B ...
#    Time_yr       (time_yr) int32 40B ...
#    RAPID_mn_cfs  (comid, time_mn) float32 1GB ...
#    RAPID_yr_cfs  (comid, time_yr) float32 107MB ...

现有代码如下,assignCoordinates函数单独运行正常,但合并时仍失败:

def assignCoordinates(df):
    df = df.assign_coords({
        "comid": df['COMID'], 
        "time_mn": fd.calcDatetimes(df, 'Time_mn', df.sizes['time_mn']), # 该函数用于计算特殊时间单位的日期时间,可正常运行
        "time_yr": fd.calcDatetimes(df, 'Time_yr', df.sizes['time_yr'])
    })
    return df

path = "path/to/files/*.nc"
ds = xr.open_mfdataset(path, preprocess=assignCoordinates, combine='by_coords', decode_times=False)

ds

怀疑预处理后的数据集未被open_mfdataset正确识别,且移除decode_times=False后会出现时间解码错误。求无需逐个打开数据集的解决方案。


最小可复现示例

生成测试文件的代码

import xarray as xr 
import numpy as np
import pandas as pd

np.random.seed(0)
temperature = 15 + 8 * np.random.randn(2, 3, 4)
precipitation = 10 * np.random.rand(2, 3, 4)
lon = [-99.83, -99.32]
lat = [42.25, 42.21]
instruments = ["manufac1", "manufac2", "manufac3"]
time = pd.date_range("2014-09-06", periods=4)
reference_time = pd.Timestamp("2014-09-05")
ds = xr.Dataset(
    data_vars=dict(
        temperature=(["loc", "instrument", "time"], temperature),
        precipitation=(["loc", "instrument", "time"], precipitation),
    ),
    attrs=dict(description="Weather related data."),
)
for i in range(1,4):
    ds.to_netcdf(f'yourdirectory/test{i}.nc') # 请修改此处路径

测试代码

import xarray as xr

def assignCoordinates(df):
    df = df.assign_coords({
        "loc": df['loc'],
        "instrument": df['instrument'],
        "time": df['time']
    })
    return df

ds = xr.open_mfdataset('yourdirectory/*.nc', preprocess=assignCoordinates, combine='by_coords') # 请修改此处路径
ds

解决方案

核心原因

combine='by_coords'要求数据集的维度坐标具备可排序/匹配的明确信息,且预处理后需确保坐标与对应维度绑定;另外decode_times=False可能导致时间坐标未被识别为可排序类型,或原数据变量与坐标同名造成混淆。

1. 修复预处理函数

确保坐标与维度绑定,同时删除冗余的原数据变量(避免同名干扰):

def assignCoordinates(df):
    # 计算时间并转为datetime64类型(确保可排序)
    time_mn = fd.calcDatetimes(df, 'Time_mn', df.sizes['time_mn']).astype('datetime64[ns]')
    time_yr = fd.calcDatetimes(df, 'Time_yr', df.sizes['time_yr']).astype('datetime64[ns]')
    
    # 分配坐标并删除原变量
    df = df.assign_coords({
        "comid": df['COMID'], 
        "time_mn": time_mn,
        "time_yr": time_yr
    }).drop(['COMID', 'Time_mn', 'Time_yr'])
    
    # 显式绑定维度与坐标,确保xarray识别
    df = df.set_index(comid='comid', time_mn='time_mn', time_yr='time_yr')
    return df

2. 调整合并参数

若明确数据集的拆分维度(如time_mn或comid),可改用combine='nested'并指定合并维度,避免自动检测失败:

ds = xr.open_mfdataset(
    path, 
    preprocess=assignCoordinates, 
    combine='nested',
    concat_dim='time_mn', # 根据实际拆分维度修改
    decode_times=False
)

3. 修复最小复现示例

原测试代码生成的数据集无对应维度的变量,导致df['loc']报错,需补充维度变量:

# 修改生成测试文件的代码,添加维度对应变量
ds = xr.Dataset(
    data_vars=dict(
        temperature=(["loc", "instrument", "time"], temperature),
        precipitation=(["loc", "instrument", "time"], precipitation),
        loc=(["loc"], [0,1]),
        instrument=(["instrument"], instruments),
        time=(["time"], time)
    ),
    attrs=dict(description="Weather related data."),
)

再运行测试代码即可正常合并。


内容的提问来源于stack exchange,提问作者MKF

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 14:57:00