You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何转换PyArrow Table以适配PyArrow.compute方法求和结构体字段

问题描述

我通过json.read_json()创建了一个PyArrow Table,其 schema 如下:

pyarrow.Table
_id: string
user_id: string
local_date_str: timestamp[s]
datetime: timestamp[s]
data: struct<aggregatable_quantity: double>
  child 0, aggregatable_quantity: double
----

需要对所有aggregatable_quantity值求和,尝试直接提取data列(返回pyarrow.ChunkedArray类型)调用pc.sum(dat)时,触发报错:

pyarrow.lib.ArrowNotImplementedError: Function 'sum' has no kernel matching input types (struct<aggregatable_quantity: double>)

原因是data列是struct类型,sum函数无法直接处理。目前通过转Pandas的方式实现需求,但步骤繁琐:

dat2 = dat.chunk(0)
datTable = pa.Table.from_struct_array(dat2)
pandArray = datTable.to_pandas()
datArray = pa.Array.from_pandas(pandArray.aggregatable_quantity)
my_sum = pc.sum(datArray)

希望找到更简洁的PyArrow原生方法,从Table中获取可用于PyArrow.compute函数的数值类数组对象。

优化解法

有两种PyArrow原生方式可以直接提取struct中的数值子列,无需转Pandas:

方法1:通过点式列名访问

PyArrow支持用表名['父列.子字段名']的方式直接提取struct中的子字段,返回对应的ChunkedArray:

import pyarrow.compute as pc

# 假设你的PyArrow Table对象为table
numeric_array = table['data.aggregatable_quantity']
my_sum = pc.sum(numeric_array)

方法2:使用struct_field函数提取

通过pyarrow.compute.struct_field函数,从struct类型的列中指定提取目标子字段:

import pyarrow as pa
import pyarrow.compute as pc

data_col = table.column('data')
numeric_array = pc.struct_field(data_col, 'aggregatable_quantity')
my_sum = pc.sum(numeric_array)

两种方法都能直接得到可用于pc.sum的数值类数组,步骤简洁且完全基于PyArrow原生实现。

内容的提问来源于stack exchange,提问作者zerzevul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 18:03:17