如何将非纯列化数据高效存储为Apache Arrow纯列格式?
如何高效将非纯列化Numpy数据转换为Apache Arrow格式的纯列化数据?
背景
此前我提过类似的非纯列化数据存储问题,但当时针对的是DuckDB。DuckDB与Apache Arrow并不相同,无法通过SQL直接向Arrow表或数组插入数据,只能间接操作。Arrow数组的构建逻辑和Numpy数组类似,我猜测或许可以借助Arrow Tensors实现需求,但不确定,因此提出此问题。
数据说明
非纯列化原始数据
"hello", "2024 JAN", "2024 FEB" "a", 0, 1
目标纯列化格式
"hello", "year", "month", "value" "a", 2024, "JAN", 0 "a", 2024, "FEB", 1
数据的Numpy数组形式
import numpy as np data = np.array( [["hello", "2024 JAN", "2024 FEB"], ["a", "0", "1"]], dtype="<U" )
输出结果:
array([['hello', '2024 JAN', '2024 FEB'], ['a', '0', '1']], dtype='<U8')
常规暴力转换方法
先在Python中将数据重组为纯列化格式,再生成Arrow数组,代码如下:
import re from typing import Dict, Tuple data_header = data[0] data_proper = data[1:] date_pattern = re.compile(r"(?P<year>[\d]+) (?P<month>JAN|FEB)") common_labels: list[str] = [] header_to_date: Dict[str, Tuple[int, str]] = dict() for header in data_header: if matches := date_pattern.match(header): year, month = int(matches["year"]), str(matches["month"]) header_to_date[header] = (year, month) else: common_labels.append(header) # 构建纯列化数组 new_rows_per_old_row = len(header_to_date) purely_columnar = np.empty( ( 1 + data_proper.shape[0] * new_rows_per_old_row, len(common_labels) + 3, ), dtype=np.object_, ) purely_columnar[0] = common_labels + ["year", "month", "value"] for rx, row in enumerate(data_proper): common_data = [] ym_data = [] for header, element in zip(data_header, row): if header in common_labels: common_data.append(element) else: year, month = header_to_date[header] ym_data.append([year, month, element]) for yx, year_month_value in enumerate(ym_data): purely_columnar[ 1 + rx * new_rows_per_old_row + yx, : len(common_labels) ] = common_data purely_columnar[ 1 + rx * new_rows_per_old_row + yx, len(common_labels) : ] = year_month_value print(f"{purely_columnar=}")
输出结果:
purely_columnar= array([[np.str_('hello'), 'year', 'month', 'value'], [np.str_('a'), 2024, 'JAN', np.str_('0')], [np.str_('a'), 2024, 'FEB', np.str_('1')]], dtype=object)
基于重组后的数组生成Arrow表:
import pyarrow as pa column_types = [pa.string(), pa.int64(), pa.string(), pa.string()] pa.table( [ pa.array(purely_columnar[1:, cx], type=column_types[cx]) for cx in range(purely_columnar.shape[1]) ], names=purely_columnar[0], )
输出结果:
pyarrow.Table hello: string year: int64 month: string value: string ---- hello: [["a","a"]] year: [[2024,2024]] month: [["JAN","FEB"]] value: [["0","1"]]
提问
除了这种先暴力转换为纯列化Numpy数组再生成Arrow表的方法,有没有更高效的方式直接将原始非纯列化数据转换为纯列化的Apache Arrow格式?
内容的提问来源于stack exchange,提问作者bzm3r
相关产品推荐
相关产品推荐

