You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将非纯列化数据高效存储为Apache Arrow纯列格式?

如何高效将非纯列化Numpy数据转换为Apache Arrow格式的纯列化数据?

背景

此前我提过类似的非纯列化数据存储问题,但当时针对的是DuckDB。DuckDB与Apache Arrow并不相同,无法通过SQL直接向Arrow表或数组插入数据,只能间接操作。Arrow数组的构建逻辑和Numpy数组类似,我猜测或许可以借助Arrow Tensors实现需求,但不确定,因此提出此问题。

数据说明

非纯列化原始数据

"hello", "2024 JAN", "2024 FEB"
"a", 0, 1

目标纯列化格式

"hello", "year", "month", "value"
"a", 2024, "JAN", 0
"a", 2024, "FEB", 1

数据的Numpy数组形式

import numpy as np

data = np.array(
    [["hello", "2024 JAN", "2024 FEB"], ["a", "0", "1"]], dtype="<U"
)

输出结果:

array([['hello', '2024 JAN', '2024 FEB'],
       ['a', '0', '1']], dtype='<U8')

常规暴力转换方法

先在Python中将数据重组为纯列化格式,再生成Arrow数组,代码如下:

import re
from typing import Dict, Tuple

data_header = data[0]
data_proper = data[1:]

date_pattern = re.compile(r"(?P<year>[\d]+) (?P<month>JAN|FEB)")
common_labels: list[str] = []
header_to_date: Dict[str, Tuple[int, str]] = dict()

for header in data_header:
    if matches := date_pattern.match(header):
        year, month = int(matches["year"]), str(matches["month"])
        header_to_date[header] = (year, month)
    else:
        common_labels.append(header)

# 构建纯列化数组
new_rows_per_old_row = len(header_to_date)
purely_columnar = np.empty(
    (
        1 + data_proper.shape[0] * new_rows_per_old_row,
        len(common_labels) + 3,
    ),
    dtype=np.object_,
)
purely_columnar[0] = common_labels + ["year", "month", "value"]

for rx, row in enumerate(data_proper):
    common_data = []
    ym_data = []
    for header, element in zip(data_header, row):
        if header in common_labels:
            common_data.append(element)
        else:
            year, month = header_to_date[header]
            ym_data.append([year, month, element])

    for yx, year_month_value in enumerate(ym_data):
        purely_columnar[
            1 + rx * new_rows_per_old_row + yx, : len(common_labels)
        ] = common_data
        purely_columnar[
            1 + rx * new_rows_per_old_row + yx, len(common_labels) :
        ] = year_month_value

print(f"{purely_columnar=}")

输出结果:

purely_columnar=
array([[np.str_('hello'), 'year', 'month', 'value'],
       [np.str_('a'), 2024, 'JAN', np.str_('0')],
       [np.str_('a'), 2024, 'FEB', np.str_('1')]], dtype=object)

基于重组后的数组生成Arrow表:

import pyarrow as pa

column_types = [pa.string(), pa.int64(), pa.string(), pa.string()]
pa.table(
    [
        pa.array(purely_columnar[1:, cx], type=column_types[cx])
        for cx in range(purely_columnar.shape[1])
    ],
    names=purely_columnar[0],
)

输出结果:

pyarrow.Table
hello: string
year: int64
month: string
value: string
----
hello: [["a","a"]]
year: [[2024,2024]]
month: [["JAN","FEB"]]
value: [["0","1"]]

提问

除了这种先暴力转换为纯列化Numpy数组再生成Arrow表的方法,有没有更高效的方式直接将原始非纯列化数据转换为纯列化的Apache Arrow格式?


内容的提问来源于stack exchange,提问作者bzm3r

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 20:33:12