You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyArrow写入Parquet文件时必填字段处理Null值异常问题

问题:PyArrow处理含Null值的非nullable字段时,写入Parquet后数据异常

测试代码

def test_pyarow():
    import pyarrow as pa
    import pyarrow.parquet
    import pandas as pd

    fields = [pa.field('id', pa.string(), nullable=False),
              pa.field('name', pa.string(), nullable=False)]
    array = [pa.array(['10', '11', '12', '13']),
             pa.array(['AAA', None, 'BBB', 'CCC'])]
    table = pa.Table.from_arrays(array, schema=pa.schema(fields))
    pyarrow.parquet.write_table(table, 'test_arrow.parquet', compression='SNAPPY', use_compliant_nested_type=True)
    df = pd.read_parquet("/Users/fki/Documents/git/Demo/bq_api/test_arrow.parquet", engine='pyarrow')
    print("\n\n\n")
    print(df)

现象对比

  • 当字段nullable=True时,输出结果正常:
id  name
0  10   AAA
1  11  None
2  12   BBB
3  13   CCC
  • 当字段nullable=False时,输出数据异常:
id name
0  10  AAA
1  11  BBB
2  12  CCC
3  13  AAA

原因与解决办法

核心原因

你定义的schema中name字段被标记为nullable=False,但传入的数组却包含None值,这直接违反了schema的非空约束。PyArrow在这种场景下不会主动抛出错误,而是会对Null值做未定义的内存处理——具体表现为跳过Null值的位置,直接读取后续有效数据填充,导致整体数据偏移,最终出现重复、错位的异常结果。

解决办法

  1. 保证数据匹配schema约束:如果字段设为nullable=False,必须提前过滤或填充数组中的Null值,比如用默认值替换:
    # 将name数组中的None替换为默认值
    name_values = ['AAA', None, 'BBB', 'CCC']
    filled_name_array = pa.array([v if v is not None else 'UNKNOWN' for v in name_values])
    
  2. 显式校验数据合法性:在创建Table前主动检查数据是否符合schema约束,发现Null值时直接抛出错误:
    schema = pa.schema(fields)
    # 遍历校验每个字段的非空约束
    for idx, field in enumerate(schema):
        if not field.nullable and array[idx].null_count > 0:
            raise ValueError(f"字段{field.name}不允许为Null,但存在{array[idx].null_count}个Null值")
    table = pa.Table.from_arrays(array, schema=schema)
    
  3. 保留nullable=True:如果业务场景允许字段为空,直接将nullable设为True,PyArrow会正确处理Null值,不会出现数据错位问题。

内容的提问来源于stack exchange,提问作者Gaurang Shah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 02:57:04