使用fastparquet引擎写入Parquet时触发ValueError的原因排查
问题
使用Pandas DataFrame的to_parquet方法并指定engine="fastparquet"生成Parquet文件时触发ValueError,换用engine="pyarrow"则正常运行。相关代码及错误信息如下:
代码片段
import datetime as dt import pandas as pd timestamp = dt.datetime.now().strftime("%Y%m%dT%H%M%S") print('timestamp:', timestamp) print('dataset_df:', dataset_df, type(dataset_df)) dataset_df.to_parquet(f"dataset_{timestamp}.parquet", engine="fastparquet")
运行输出
timestamp: 20230213T163544 dataset_df: [7001 rows x 51 columns] <class 'pandas.core.frame.DataFrame'>
错误栈
Traceback (most recent call last): File "c:\users\projects\prepro.py", line 320, in prepro dataset_df.to_parquet( File "C:\Users\user_name\.conda\envs\py38\lib\site-packages\pandas\util\_decorators.py", line 211, in wrapper return func(*args, **kwargs) File "C:\Users\.conda\envs\py38\lib\site-packages\pandas\core\frame.py", line 2975, in to_parquet return to_parquet( File "C:\Users\.conda\envs\py38\lib\site-packages\pandas\io\parquet.py", line 430, in to_parquet impl.write( se, type = find_type(data[column], fixed_text=fixed, File "C:\Users\.conda\envs\py38\lib\site-packages\fastparquet\writer.py", line 126, in find_type object_encoding = infer_object_encoding(data) File "C:\Users\.conda\envs\py38\lib\site-packages\fastparquet\writer.py", line 343, in infer_object_encoding if pd.isna(i): ValueError: The truth value of an array with more than one element is ambiguous. Use a.any() or a.all()
错误成因分析
这个错误的核心原因是你的DataFrame中存在元素为数组/列表的列,fastparquet与pyarrow对这类列的处理逻辑存在差异:
- fastparquet的
infer_object_encoding函数会遍历列中的每个元素,当遇到数组/列表类型的元素时,pd.isna(i)会返回一个布尔数组而非单个布尔值,直接用这个数组做if判断就会触发"真值判断模糊"的错误——Python无法确定你要判断数组中所有元素为真还是至少一个为真。 - pyarrow对包含数组/列表的嵌套结构列有更完善的处理逻辑,能正确识别并写入Parquet文件,因此不会报错。
内容的提问来源于stack exchange,提问作者tw0930
相关产品推荐
相关产品推荐

