You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyArrow如何转换ListArray内Struct schema合并新旧Parquet数据

解决方案

出现该报错的原因是PyArrow当前对多层嵌套结构的全表cast支持不完善,嵌套在List或上层Struct内的字段类型变更无法通过整表cast()方法直接处理。你可以采用读Parquet时直接指定目标schema的方式完成类型对齐,该方案对嵌套结构基础类型的向上转换(int转float)兼容性远高于读后cast。

最简实现步骤

  1. 以新表(active表)的schema作为统一的目标schema
  2. 读取新旧两张表时都指定该目标schema,读取阶段会自动完成旧表int字段到float字段的安全转换
  3. 读取完成后两张表schema完全一致,可直接合并

代码示例

import pyarrow as pa
import pyarrow.parquet as pq

# 直接读取新表的schema作为统一目标schema,无需手动构造全量字段
target_schema = pq.read_schema("path/to/active/parquet")

# 读表时传入schema参数,自动完成类型对齐
active = pq.read_table("path/to/active/parquet", schema=target_schema)
hist = pq.read_table("path/to/hist/parquet", schema=target_schema)

# 直接合并即可
combined = pa.concat_tables([active, hist])

# 写入合并后的文件
pq.write_table(combined, "path/to/output/combined.parquet")

注意事项

  • int64转float64属于安全向上转换,只要你的业务数值不超过2^53,不会出现精度丢失问题
  • 如果需要自定义部分字段的类型,也可以手动构造pa.Schema对象作为目标schema传入读方法,兼容所有嵌套结构的类型调整

内容的提问来源于stack exchange,提问作者matthewmturner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 23:36:05