使用Geopandas Pyogrio引擎读取GeoPackage时出现时间戳转换错误
解决Pyogrio读取GeoPackage时间戳列时的ArrowInvalid错误
问题背景:读取含30万条记录的GeoPackage时,Fiona引擎速度过慢,改用Pyogrio引擎(Arrow模式)后速度大幅提升,但读取包含时间戳列时触发ArrowInvalid: Casting from timestamp[ms] to timestamp[ns] would result in out of bounds timestamp: -59103216000000错误。原因是该毫秒级时间戳转纳秒后,超出了Arrow的timestamp[ns]支持范围(约1677-09-21至2262-04-11),且Pyogrio暂不支持读取前指定数据类型,以下是可行解决方案:
方案1:拆分读取,单独处理时间戳列
先读取不含时间戳的列,再单独读取时间戳列并转换为合适格式,最后合并数据:
import geopandas as gp import pyogrio import pandas as pd path = "mnt2/Base.gpkg" # 先读取不含时间戳的核心列 columns_without_ts = ['BUILDING_ID','GEOMETRY'] geb = gp.read_file(path, layer='BUILDING', engine='pyogrio', columns=columns_without_ts, use_arrow=True) # 单独读取时间戳列,用read_basic避免Arrow自动转换逻辑 ts_df = pyogrio.read_basic(path, layer='BUILDING', columns=['你的时间戳列名']) # 将毫秒级时间戳直接转为pandas支持的datetime64[ms]类型 ts_df['你的时间戳列名'] = pd.to_datetime(ts_df['你的时间戳列名'], unit='ms') # 合并到GeoDataFrame geb = geb.join(ts_df)
方案2:禁用Arrow引擎读取
关闭Pyogrio的Arrow模式,使用传统读取方式,避开Arrow的时间戳转换逻辑,速度虽略低于Arrow模式,但仍远快于Fiona:
import geopandas as gp path = "mnt2/Base.gpkg" columns_to_read = ['BUILDING_ID','GEOMETRY','你的时间戳列名'] # 去掉use_arrow=True参数,默认使用非Arrow模式读取 geb = gp.read_file(path, layer='BUILDING', engine='pyogrio', columns=columns_to_read)
方案3:直接读取Arrow表并手动转换类型
用Pyogrio直接读取为Arrow表,手动处理时间戳列的类型转换后,再转为GeoDataFrame:
import geopandas as gp import pyogrio from pyogrio.arrow import to_geodataframe path = "mnt2/Base.gpkg" columns_to_read = ['BUILDING_ID','GEOMETRY','你的时间戳列名'] # 读取为Arrow表 arrow_table = pyogrio.read_arrow(path, layer='BUILDING', columns=columns_to_read) # 将时间戳列强制保留为timestamp[ms]类型,再转为pandas的datetime ts_column = arrow_table['你的时间戳列名'].cast('timestamp[ms]').to_pandas() # 替换Arrow表中的原时间戳列 arrow_table = arrow_table.set_column(arrow_table.schema.get_field_index('你的时间戳列名'), '你的时间戳列名', ts_column) # 转为GeoDataFrame geb = to_geodataframe(arrow_table)
方案4:预处理源数据(若有权限)
如果可以修改源GeoPackage,将时间戳列转换为字符串类型,或者调整时间范围到Arrow的timestamp[ns]支持区间内,之后再用Pyogrio读取。
内容的提问来源于stack exchange,提问作者Rarts
相关产品推荐
相关产品推荐

