You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas调用read_feather报ArrowInvalid timestamp[us]转ns越界错误如何解决

问题原因

该报错是由于Pandas 1.2.1版本读写Feather文件时,默认强制将所有时间戳转换为纳秒(ns)精度存储,但你从SQL读取的数据集里存在超出纳秒时间戳表示范围的时间值(早于1677-09-21或晚于2262-04-11),转换时触发越界异常。

可行解决方案
  • 方案1:升级Pandas版本至1.3.0及以上
    1.3.0之后的Pandas版本优化了Feather的读写逻辑,支持保留微秒(us)精度的时间戳,无需强制转纳秒,是解决该问题成本最低的方案。升级后原有读写代码无需修改即可正常运行。
  • 方案2:导出前转换时间列类型(适配无法升级版本的场景)
    先提取所有datetime类型的列,统一转换为字符串或date类型,避开精度转换逻辑,示例代码:
    # 提取所有时间列
    time_cols = df.select_dtypes(include=['datetime64']).columns
    # 转换为字符串格式
    df[time_cols] = df[time_cols].astype(str)
    # 若不需要时间精度,也可转换为date类型
    # df[time_cols] = df[time_cols].dt.date
    
    # 再执行导出操作
    df.to_feather("some_file.feather")
    
    读取后如果需要使用时间类型,再执行对应转换即可。
  • 方案3:读写时显式指定时间精度
    导出时指定时间列存储精度为微秒,读取时也显式声明精度,避免强制转纳秒,示例代码:
    # 导出时指定精度
    time_cols = df.select_dtypes(include=['datetime64']).columns
    dtype_dict = {col: "timestamp[us]" for col in time_cols}
    df.to_feather("some_file.feather", dtype=dtype_dict)
    
    # 读取时匹配精度
    pd.read_feather("some_file.feather", dtype=dtype_dict)
    
  • 方案4:清洗越界时间值
    若超出范围的时间值为无效脏数据,可直接过滤或替换为NaT空值,示例代码:
    time_cols = df.select_dtypes(include=['datetime64']).columns
    # 替换超出纳秒表示范围的时间为NaT
    df[time_cols] = df[time_cols].apply(lambda x: x.where(x >= pd.Timestamp.min, pd.NaT))
    

内容的提问来源于stack exchange,提问作者Adrien Pacifico

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 00:21:01