PyArrow:如何指定部分Schema并自动推断动态列类型?
解决PyArrow指定已知列类型并自动推断动态列类型的问题
要实现指定已知列类型、同时让动态列自动推断类型的需求,不能直接传入仅包含已知列的schema——因为PyArrow会严格按照传入的schema过滤字段,未定义的列会被直接丢弃。正确的做法是先让PyArrow自动生成所有列的临时schema,再替换已知列的指定类型,具体实现如下:
实现步骤
- 第一步:基于输入数据自动生成包含所有列的临时schema
- 第二步:替换临时schema中已知列的类型为指定类型
- 第三步:使用修改后的schema创建表
示例代码
import pyarrow as pa # 准备数据 n_legs = pa.array([2, 4, 5, 100]) animals = pa.array(["Flamingo", "Horse", "Brittle stars", "Centipede"]) pydict = {'n_legs': n_legs, 'animals': animals} # 1. 自动生成包含所有列的临时schema temp_schema = pa.Schema.from_pydict(pydict) # 2. 定义需要指定类型的已知列 specified_types = {'n_legs': pa.int32()} # 3. 构建最终schema:替换已知列类型,其余保留自动推断结果 final_fields = [] for field in temp_schema.fields: if field.name in specified_types: final_fields.append(pa.field(field.name, specified_types[field.name])) else: final_fields.append(field) final_schema = pa.schema(final_fields) # 4. 创建表 table = pa.Table.from_pydict(pydict, schema=final_schema) print(table)
运行结果
pyarrow.Table n_legs: int32 animals: string ---- n_legs: [[2,4,5,100]] animals: [["Flamingo","Horse","Brittle stars","Centipede"]]
可以看到,n_legs列被指定为int32类型,animals列则自动推断为string类型,所有列都被保留。
内容的提问来源于stack exchange,提问作者Brian
相关产品推荐
相关产品推荐

