向Arrow文件追加表时PyArrow字典类型报错的解决方案咨询
问题描述
处理大型Pandas DataFrame数据集时,需分块读取并追加写入PyArrow Feather文件。为压缩重复率高的字符串字段(如全量重复的symbol、仅约20种取值的exch等),在Schema中使用了pa.dictionary()类型,但运行时报错:
"Dictionary replacement detected when writing IPC file format. Arrow IPC files only support a single non-delta dictionary for a given field across all batches."
当前使用的代码如下:
import pyarrow as pa fp = pa.OSFile('my_file.arrow', 'wb') schema = pa.schema( [ ('localtime', pa.timestamp('ns')), ('exchtime', pa.timestamp('ns')), ('msgtype', pa.string()), ('symbol', pa.dictionary(pa.int8(), pa.string())), ('exch', pa.string()), ('price', pa.float64()), ('size', pa.float64()), ('side', pa.dictionary(pa.int8(), pa.string())), ('ref', pa.uint64()), ('oldref', pa.uint64()), ('mpid', pa.string()), ('esd', pa.dictionary(pa.int8(), pa.string())), ('platform', pa.dictionary(pa.int8(), pa.string())), ] ) writer = pa.ipc.new_file(fp, schema) # ... 分块读取数据的逻辑 ... for df in [... list of dataframes chunks ...]: table = pa.Table.from_pandas(df, schema) writer.write(table) writer.close() fp.close()
已尝试操作:
- 不使用字典类型时,代码正常运行,但文件体积过大
- 某列所有数据块取值相同时,字典类型可正常运行
- 为每个数据块推断Schema而非显式指定时,因Schema不统一无法运行
疑问:若提前知晓字段的所有可能字符串取值,是否可显式指定字符串到整数的字典映射?
解决方案
可以提前显式指定字典映射,确保所有分块数据使用统一的字典编码,从根源避免字典替换的问题,具体实现如下:
1. 预定义字段的字典映射
针对需要字典编码的字段,提前创建字符串到整数的映射表,并生成带固定字典的PyArrow类型:
# 预定义各字段的完整取值与整数映射 symbol_mapping = {'AAPL': 0, 'MSFT': 1, 'GOOG': 2} side_mapping = {'BUY': 0, 'SELL': 1} esd_mapping = {'YES': 0, 'NO': 1} platform_mapping = {'WEB': 0, 'APP': 1, 'API': 2} # 创建带固定字典的PyArrow字典类型 symbol_type = pa.dictionary( pa.int8(), pa.string(), ordered=False, dictionary=pa.array(list(symbol_mapping.keys())) ) side_type = pa.dictionary( pa.int8(), pa.string(), ordered=False, dictionary=pa.array(list(side_mapping.keys())) ) esd_type = pa.dictionary( pa.int8(), pa.string(), ordered=False, dictionary=pa.array(list(esd_mapping.keys())) ) platform_type = pa.dictionary( pa.int8(), pa.string(), ordered=False, dictionary=pa.array(list(platform_mapping.keys())) )
2. 更新Schema定义
用预定义的带固定字典的类型替换原Schema中的普通字典类型:
schema = pa.schema( [ ('localtime', pa.timestamp('ns')), ('exchtime', pa.timestamp('ns')), ('msgtype', pa.string()), ('symbol', symbol_type), ('exch', pa.string()), # 若exch取值固定,也可改为字典类型 ('price', pa.float64()), ('size', pa.float64()), ('side', side_type), ('ref', pa.uint64()), ('oldref', pa.uint64()), ('mpid', pa.string()), ('esd', esd_type), ('platform', platform_type), ] )
3. 分块转换时强制使用统一字典
在将每个DataFrame分块转为PyArrow Table前,先把字符串列映射为对应的整数,确保编码一致:
fp = pa.OSFile('my_file.arrow', 'wb') writer = pa.ipc.new_file(fp, schema) for df in [... list of dataframes chunks ...]: # 将字符串列映射为预定义的整数 df['symbol'] = df['symbol'].map(symbol_mapping) df['side'] = df['side'].map(side_mapping) df['esd'] = df['esd'].map(esd_mapping) df['platform'] = df['platform'].map(platform_mapping) # 转换为Table时指定Schema,确保类型匹配 table = pa.Table.from_pandas(df, schema=schema, preserve_index=False) writer.write(table) writer.close() fp.close()
关键注意事项
- 必须确保所有分块数据中的字符串取值都包含在预定义的字典中,否则会出现映射失败或编码错误
- 若后续有新增取值,需提前更新字典映射表和对应的PyArrow类型,否则无法写入新值
- 对于
exch这类取值较少的字段,同样可以套用上述逻辑改为字典类型,进一步压缩文件体积
内容的提问来源于stack exchange,提问作者Patrick Nichols
相关产品推荐
相关产品推荐

