Pandas多列索引DataFrame存Parquet追加数据报错求助
问题:追加Parquet文件后加载时出现分类字典不匹配错误
复现代码
初始数据写入与加载
import pandas as pd # 构造初始数据集 initial_df = pd.DataFrame([ {'a': 0, 'b': 0, 'c': 10.898}, {'a': 0, 'b': 1, 'c': 1.88}, {'a': 1, 'b': 0, 'c': 108.1}, {'a': 1, 'b': 1, 'c': 10.898} ]) initial_df.set_index(['a', 'b'], inplace=True) # 写入Parquet文件 initial_df.to_parquet('test.parquet', engine='fastparquet', compression='GZIP', append=False, index=True) # 正常加载文件 read_df = pd.read_parquet('test.parquet', engine='fastparquet')
追加数据与报错场景
# 构造追加数据集 additional_df = pd.DataFrame([ {'a': 2, 'b': 0, 'c': 10.898}, {'a': 2, 'b': 1, 'c': 1.88}, {'a': 3, 'b': 0, 'c': 108.1}, {'a': 3, 'b': 1, 'c': 10.898} ]) additional_df.set_index(['a', 'b'], inplace=True) # 执行追加操作 additional_df.to_parquet('test.parquet', engine='fastparquet', compression='GZIP', append=True, index=True) # 加载时触发错误 read_df = pd.read_parquet('test.parquet', engine='fastparquet')
错误信息
RuntimeError: Different dictionaries encountered while building categorical
错误位置:pandas\io\parquet.py:358
版本信息
- python: 3.10.8
- pandas: 1.5.1
- fastparquet: 0.8.3(测试过0.5.0版本,问题依旧)
已尝试的调试动作
- 跟踪fastparquet源码发现
fastparquet\core.py:170中的read_col函数被多次调用,索引被重复写入,第二次写入触发报错 - 调整
read_parquet的index参数,未解决问题
解决方案
方案1:追加时不写入索引,读取后重建索引
将索引转为普通列写入,追加完成后读取再重新设置索引,规避索引的分类字典冲突:
# 初始数据处理:重置索引为列,写入时不存索引 initial_df.reset_index(inplace=True) initial_df.to_parquet('test.parquet', engine='fastparquet', compression='GZIP', append=False, index=False) # 追加数据同样处理 additional_df.reset_index(inplace=True) additional_df.to_parquet('test.parquet', engine='fastparquet', compression='GZIP', append=True, index=False) # 读取后重建索引 read_df = pd.read_parquet('test.parquet', engine='fastparquet') read_df.set_index(['a', 'b'], inplace=True)
方案2:升级fastparquet到最新稳定版本
旧版本fastparquet在处理带索引的Parquet追加时存在bug,升级到0.10.x及以上版本(如0.12.0)可修复该问题,升级后直接使用原代码即可正常追加和读取。
方案3:切换为pyarrow引擎
pyarrow在Parquet追加和索引兼容性上表现更稳定,切换引擎后无需大幅修改代码:
# 初始写入 initial_df.to_parquet('test.parquet', engine='pyarrow', compression='GZIP', append=False, index=True) # 追加数据 additional_df.to_parquet('test.parquet', engine='pyarrow', compression='GZIP', append=True, index=True) # 正常加载 read_df = pd.read_parquet('test.parquet', engine='pyarrow')
内容的提问来源于stack exchange,提问作者KerikoN
相关产品推荐
相关产品推荐

