使用fastparquet向ADLS2文件追加DataFrame数据报错求助
问题场景
在ADLS2中已有一个Parquet文件,执行以下代码尝试追加数据失败:
filepath = "abfs://shopifyparquet/test/parquet/LIVE/filename" adls2_data_df.to_parquet(path=filepath,engine='fastparquet',storage_options={'account_name': 'test', 'account_key': 'mykey'},append=True)
本地文件测试时fastparquet可正常追加,但ADLS2上执行抛出异常:
File mode not supported
Exception ignored in: <function AzureBlobFile.del at 0x000001CBA3BC6950>
self.close()
File "C:\Users\Sivasankar.Muthuraju\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\adlfs\spec.py", line 1851, in close
super().close()
File "C:\Users\Sivasankar.Muthuraju\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\fsspec\spec.py", line 1740, in close
if not self.forced:
AttributeError: 'AzureBlobFile' object has no attribute 'forced'
可行解决方案
方案1:改用pyarrow引擎替代fastparquet
fastparquet在ADLS2的append模式支持存在兼容性问题,pyarrow对云存储的追加操作支持更完善。修改代码如下:
filepath = "abfs://shopifyparquet/test/parquet/LIVE/filename" adls2_data_df.to_parquet( path=filepath, engine='pyarrow', storage_options={'account_name': 'test', 'account_key': 'mykey'}, append=True )
方案2:手动合并文件(若必须使用fastparquet)
如果依赖fastparquet,可先读取原有文件,合并新数据后重新写入:
import pandas as pd from fastparquet import ParquetFile filepath = "abfs://shopifyparquet/test/parquet/LIVE/filename" storage_options={'account_name': 'test', 'account_key': 'mykey'} # 读取原有数据 pf = ParquetFile(filepath, storage_options=storage_options) existing_df = pf.to_pandas() # 合并新数据 combined_df = pd.concat([existing_df, adls2_data_df], ignore_index=True) # 重新写入覆盖原有文件 combined_df.to_parquet( path=filepath, engine='fastparquet', storage_options=storage_options )
方案3:升级adlfs和fsspec库
报错中的AttributeError是由于adlfs与fsspec版本不兼容导致的,升级到最新版本可修复该问题:
pip install --upgrade adlfs fsspec fastparquet
内容的提问来源于stack exchange,提问作者Mathuraju Sivasankar

