使用pyarrow.csv读取含Unix时间戳的CSV时类型转换报错
解决PyArrow读取CSV时Unix时间整数转Timestamp的报错问题
问题原因
PyArrow通过column_types指定pa.timestamp('s')时,会尝试直接将CSV中的字符串格式数字解析为时间戳,但它默认不支持这种转换逻辑,因此抛出ArrowInvalid错误。
解决方案
方案1:先读为整数,再转换时间戳
先将timestamp列读取为int64类型,再通过类型转换转为秒级时间戳:
import pyarrow as pa import pyarrow.csv as pv # 读取时指定列为int64 table = pv.read_csv("file.csv", convert_options=pv.ConvertOptions( column_types={'timestamp': pa.int64()} )) # 将整数列转为timestamp[s]类型 timestamp_idx = table.column_names.index('timestamp') table = table.set_column( timestamp_idx, 'timestamp', table['timestamp'].cast(pa.timestamp('s')) )
方案2:使用自定义转换器直接解析
通过converters参数自定义转换逻辑,将字符串形式的Unix时间整数直接转为时间戳:
import pyarrow as pa import pyarrow.csv as pv def parse_unix_ts(s): return pa.scalar(int(s), type=pa.timestamp('s')) convert_opts = pv.ConvertOptions( column_types={'timestamp': pa.timestamp('s')}, converters={'timestamp': parse_unix_ts} ) table = pv.read_csv("file.csv", convert_options=convert_opts)
说明
- 方案1操作简单,适合大多数常规场景;
- 方案2在读取阶段完成转换,无需额外的列修改步骤,适合对流程精简有要求的场景。
内容的提问来源于stack exchange,提问作者David Davó
相关产品推荐
相关产品推荐

