You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

向Arrow文件追加表时PyArrow字典类型报错的解决方案咨询

问题描述

处理大型Pandas DataFrame数据集时,需分块读取并追加写入PyArrow Feather文件。为压缩重复率高的字符串字段(如全量重复的symbol、仅约20种取值的exch等),在Schema中使用了pa.dictionary()类型,但运行时报错:

"Dictionary replacement detected when writing IPC file format. Arrow IPC files only support a single non-delta dictionary for a given field across all batches."

当前使用的代码如下:

import pyarrow as pa
fp = pa.OSFile('my_file.arrow', 'wb')
schema = pa.schema(
    [
        ('localtime', pa.timestamp('ns')),
        ('exchtime', pa.timestamp('ns')),
        ('msgtype', pa.string()),
        ('symbol', pa.dictionary(pa.int8(), pa.string())),
        ('exch', pa.string()),
        ('price', pa.float64()),
        ('size', pa.float64()),
        ('side', pa.dictionary(pa.int8(), pa.string())),
        ('ref', pa.uint64()),
        ('oldref', pa.uint64()),
        ('mpid', pa.string()),
        ('esd', pa.dictionary(pa.int8(), pa.string())),
        ('platform', pa.dictionary(pa.int8(), pa.string())),
    ]
)
writer = pa.ipc.new_file(fp, schema)

# ... 分块读取数据的逻辑 ...

for df in [... list of dataframes chunks ...]:
  table = pa.Table.from_pandas(df, schema)
  writer.write(table)

writer.close()
fp.close()

已尝试操作:

  • 不使用字典类型时,代码正常运行,但文件体积过大
  • 某列所有数据块取值相同时,字典类型可正常运行
  • 为每个数据块推断Schema而非显式指定时,因Schema不统一无法运行

疑问:若提前知晓字段的所有可能字符串取值,是否可显式指定字符串到整数的字典映射?


解决方案

可以提前显式指定字典映射,确保所有分块数据使用统一的字典编码,从根源避免字典替换的问题,具体实现如下:

1. 预定义字段的字典映射

针对需要字典编码的字段,提前创建字符串到整数的映射表,并生成带固定字典的PyArrow类型:

# 预定义各字段的完整取值与整数映射
symbol_mapping = {'AAPL': 0, 'MSFT': 1, 'GOOG': 2}
side_mapping = {'BUY': 0, 'SELL': 1}
esd_mapping = {'YES': 0, 'NO': 1}
platform_mapping = {'WEB': 0, 'APP': 1, 'API': 2}

# 创建带固定字典的PyArrow字典类型
symbol_type = pa.dictionary(
    pa.int8(), 
    pa.string(), 
    ordered=False, 
    dictionary=pa.array(list(symbol_mapping.keys()))
)
side_type = pa.dictionary(
    pa.int8(), 
    pa.string(), 
    ordered=False, 
    dictionary=pa.array(list(side_mapping.keys()))
)
esd_type = pa.dictionary(
    pa.int8(), 
    pa.string(), 
    ordered=False, 
    dictionary=pa.array(list(esd_mapping.keys()))
)
platform_type = pa.dictionary(
    pa.int8(), 
    pa.string(), 
    ordered=False, 
    dictionary=pa.array(list(platform_mapping.keys()))
)

2. 更新Schema定义

用预定义的带固定字典的类型替换原Schema中的普通字典类型:

schema = pa.schema(
    [
        ('localtime', pa.timestamp('ns')),
        ('exchtime', pa.timestamp('ns')),
        ('msgtype', pa.string()),
        ('symbol', symbol_type),
        ('exch', pa.string()),  # 若exch取值固定,也可改为字典类型
        ('price', pa.float64()),
        ('size', pa.float64()),
        ('side', side_type),
        ('ref', pa.uint64()),
        ('oldref', pa.uint64()),
        ('mpid', pa.string()),
        ('esd', esd_type),
        ('platform', platform_type),
    ]
)

3. 分块转换时强制使用统一字典

在将每个DataFrame分块转为PyArrow Table前,先把字符串列映射为对应的整数,确保编码一致:

fp = pa.OSFile('my_file.arrow', 'wb')
writer = pa.ipc.new_file(fp, schema)

for df in [... list of dataframes chunks ...]:
    # 将字符串列映射为预定义的整数
    df['symbol'] = df['symbol'].map(symbol_mapping)
    df['side'] = df['side'].map(side_mapping)
    df['esd'] = df['esd'].map(esd_mapping)
    df['platform'] = df['platform'].map(platform_mapping)
    
    # 转换为Table时指定Schema,确保类型匹配
    table = pa.Table.from_pandas(df, schema=schema, preserve_index=False)
    writer.write(table)

writer.close()
fp.close()

关键注意事项

  • 必须确保所有分块数据中的字符串取值都包含在预定义的字典中,否则会出现映射失败或编码错误
  • 若后续有新增取值,需提前更新字典映射表和对应的PyArrow类型,否则无法写入新值
  • 对于exch这类取值较少的字段,同样可以套用上述逻辑改为字典类型,进一步压缩文件体积

内容的提问来源于stack exchange,提问作者Patrick Nichols

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 03:10:37