You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pyArrow将pandas.DataFrame转Parquet时遇ArrowInvalid错误求助

Fixing ArrowInvalid When Converting Pandas DataFrame to Parquet with PyArrow

Great job narrowing down the issue to that object-type column with leading zeros—that's exactly where the problem lies! Let me break this down and walk you through the fixes.

Why This Error Happens

PyArrow automatically infers the schema of your DataFrame when converting it to Parquet. For columns with object dtype, if most values look like integers (even if stored as strings), PyArrow will try to cast the column to an Int64 type. But when it hits strings with leading zeros (like "00123"), it can’t convert those string values to integers, hence the ArrowInvalid error.

Solutions

Pick the approach that matches your use case:

1. Keep the Column as String Type (Preserve Leading Zeros)

If those leading zeros are meaningful (e.g., IDs, zip codes), you need to explicitly tell PyArrow to treat the column as a string instead of letting it guess. Here are two ways to do this:

  • Cast the column to string dtype first:

    import pandas as pd
    import pyarrow as pa
    import pyarrow.parquet as pq
    
    # Assume your DataFrame is named df, and the problem column is 'leading_zero_col'
    df['leading_zero_col'] = df['leading_zero_col'].astype(str)
    
    # Convert and write to Parquet
    table = pa.Table.from_pandas(df)
    pq.write_table(table, 'output.parquet')
    
  • Define a custom schema explicitly:
    Use this if you want full control over all column types:

    # Define schema with the problem column set to string
    custom_schema = pa.schema([
        ('leading_zero_col', pa.string()),
        ('numeric_col', pa.int64()),
        ('text_col', pa.string()),
        # Add all remaining columns with their correct types
    ])
    
    table = pa.Table.from_pandas(df, schema=custom_schema)
    pq.write_table(table, 'output.parquet')
    

2. Convert the Column to Integer Type (Lose Leading Zeros)

If leading zeros are just formatting and you need the values as integers, convert the column to a numeric type first. Note: leading zeros will be lost (integers don’t store formatting):

# Convert to integers (raises error if non-numeric values exist)
df['leading_zero_col'] = pd.to_numeric(df['leading_zero_col'], errors='raise')

# Or handle errors gracefully (replace invalid values with NaN)
# df['leading_zero_col'] = pd.to_numeric(df['leading_zero_col'], errors='coerce')

# Write to Parquet without issues
table = pa.Table.from_pandas(df)
pq.write_table(table, 'output.parquet')

Quick Tip

Always double-check dtypes for columns with mixed-looking values (like string-formatted numbers) before converting to Parquet. Explicitly setting types avoids PyArrow’s automatic inference causing unexpected errors.

内容的提问来源于stack exchange,提问作者Carlos P Ceballos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:02:57