You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyArrow能否对Parquet中的struct、list类型字段执行过滤操作?

PyArrow Parquet数组类型字段过滤方案

你之前的代码存在两处问题:

  • 字段名拼写错误:你的表中列名为regions(复数),代码中写的是region(单数)
  • 过滤逻辑错误:isin语法是判断左侧元素是否存在于右侧集合中,你需要的是判断数组列是否包含指定标量,应该使用列表类型专属的包含判断逻辑

正确实现代码

import pyarrow.dataset as ds

dataset = ds.dataset("./example.parquet", format="parquet")
# 过滤regions数组包含us的行,过滤逻辑会下推到Parquet扫描层执行
filtered_table = dataset.to_table(filter=ds.field("regions").list_contains("us"))

# 验证输出
print(filtered_table.to_pandas())

如果你使用的是较旧版本的PyArrow,无法直接在ds.field后调用list_contains方法,可以显式调用compute模块的对应函数:

import pyarrow.dataset as ds
from pyarrow import compute as pc

dataset = ds.dataset("./example.parquet", format="parquet")
filtered_table = dataset.to_table(filter=pc.list_contains(ds.field("regions"), "us"))

执行后输出的结果仅会保留regions数组包含us的行,符合预期。

内容的提问来源于stack exchange,提问作者jstrong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 06:36:01