You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

保存含加密列的DataFrame至Parquet时遇TypeError问题求助

解决含加密列的DataFrame保存为Parquet时的TypeError问题

问题场景

执行以下代码保存含加密列的DataFrame为Parquet文件时触发错误:

df.to_parquet(path_to_data + 'data_temp.gzip', compression='GZIP')

错误信息

TypeError                                 Traceback (most recent call last)
TypeError: Expected unicode, got quoted_name
Exception ignored in: 'fastparquet.cencoding.write_list'
Traceback (most recent call last):
  File "/manifests/venv/lib/python3.10/site-packages/fastparquet/writer.py", line 1499, in write_thrift
    return f.write(obj.to_bytes())
TypeError: Expected unicode, got quoted_name

注:该代码在无加密列的DataFrame上可正常运行。

解决方法

  • 检查并转换加密列数据类型
    加密后的数据可能是自定义类型或非标准字节类型,fastparquet无法直接序列化。先查看列类型:

    print(df[['加密列名1', '加密列名2']].dtypes)
    

    转换为标准可序列化类型:

    # 若为bytes类型,转为字符串
    df['加密列名1'] = df['加密列名1'].apply(lambda x: x.decode('utf-8') if isinstance(x, bytes) else str(x))
    # 或直接保留bytes类型(部分Parquet引擎支持)
    df['加密列名1'] = df['加密列名1'].astype('bytes')
    
  • 切换Parquet引擎至pyarrow
    fastparquet对特殊类型兼容性有限,换用pyarrow引擎试试:

    df.to_parquet(path_to_data + 'data_temp.gzip', compression='GZIP', engine='pyarrow')
    

    未安装pyarrow需先执行:pip install pyarrow

  • 修复列名的编码问题
    错误中的quoted_name可能指向列名含非Unicode特殊字符,重命名列:

    # 统一列名编码
    df.columns = [str(col) for col in df.columns]
    # 或直接重命名加密列
    df.rename(columns={'原加密列名1': 'encrypted_col1', '原加密列名2': 'encrypted_col2'}, inplace=True)
    
  • 序列化自定义加密对象
    若加密列是自定义加密对象,先序列化为JSON字符串或字节流:

    import json
    df['加密列名1'] = df['加密列名1'].apply(lambda x: json.dumps(x.__dict__))
    

内容的提问来源于stack exchange,提问作者bravopapa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 18:18:26