如何从大型PyArrow Dataset中提取唯一值?
提取PyArrow Dataset中列的唯一值(无需加载全量数据)
PyArrow目前没有直接的内置API可以一次性从Dataset中提取单列或多列的唯一值,但可以利用Dataset的扫描能力结合PyArrow的计算功能高效实现,比手动逐行处理更简洁:
单列唯一值提取
通过扫描器仅加载目标列,在批次内先去重再合并到全局集合:
import pyarrow as pa import pyarrow.dataset as ds # 加载Hive分区的Parquet数据集 dataset = ds.dataset("你的数据路径", format="parquet", partitioning="hive") # 创建仅扫描目标列的扫描器 scanner = dataset.scanner(columns=["first_name"]) distinct_values = set() for batch in scanner.to_batches(): # 对当前批次的列计算去重,转为Python列表后合并到全局集合 batch_distinct = pa.compute.distinct(batch.column("first_name")).to_pylist() distinct_values.update(batch_distinct) # 最终得到单列的所有唯一值 print(distinct_values)
多列唯一元组提取
针对多列组合,将批次中的列值配对为元组后,合并到全局集合:
scanner = dataset.scanner(columns=["first_name", "last_name"]) distinct_tuples = set() for batch in scanner.to_batches(): # 将两列的值配对为元组列表 name_pairs = list(zip( batch.column("first_name").to_pylist(), batch.column("last_name").to_pylist() )) distinct_tuples.update(name_pairs) # 最终得到多列组合的唯一元组 print(distinct_tuples)
封装成复用函数
如果需要多次使用,可以封装成函数,同时兼容单列和多列场景:
def get_distinct_from_dataset(dataset, columns): scanner = dataset.scanner(columns=columns) distinct_set = set() for batch in scanner.to_batches(): if len(columns) == 1: # 单列场景用PyArrow内置去重减少数据量 batch_vals = pa.compute.distinct(batch.column(columns[0])).to_pylist() distinct_set.update(batch_vals) else: # 多列场景转为元组集合 batch_vals = list(zip(*[batch.column(c).to_pylist() for c in columns])) distinct_set.update(batch_vals) return distinct_set # 使用示例 single_col_distinct = get_distinct_from_dataset(dataset, ["first_name"]) multi_col_distinct = get_distinct_from_dataset(dataset, ["first_name", "last_name"])
这种方式仅加载需要的列,且在批次内先做去重,能有效降低内存占用,同时比手动逐行处理更高效。
内容的提问来源于stack exchange,提问作者Gabriele Giuseppini
相关产品推荐
相关产品推荐

