使用pyArrow将pandas.DataFrame转Parquet时遇ArrowInvalid错误求助
Great job narrowing down the issue to that object-type column with leading zeros—that's exactly where the problem lies! Let me break this down and walk you through the fixes.
Why This Error Happens
PyArrow automatically infers the schema of your DataFrame when converting it to Parquet. For columns with object dtype, if most values look like integers (even if stored as strings), PyArrow will try to cast the column to an Int64 type. But when it hits strings with leading zeros (like "00123"), it can’t convert those string values to integers, hence the ArrowInvalid error.
Solutions
Pick the approach that matches your use case:
1. Keep the Column as String Type (Preserve Leading Zeros)
If those leading zeros are meaningful (e.g., IDs, zip codes), you need to explicitly tell PyArrow to treat the column as a string instead of letting it guess. Here are two ways to do this:
Cast the column to string dtype first:
import pandas as pd import pyarrow as pa import pyarrow.parquet as pq # Assume your DataFrame is named df, and the problem column is 'leading_zero_col' df['leading_zero_col'] = df['leading_zero_col'].astype(str) # Convert and write to Parquet table = pa.Table.from_pandas(df) pq.write_table(table, 'output.parquet')Define a custom schema explicitly:
Use this if you want full control over all column types:# Define schema with the problem column set to string custom_schema = pa.schema([ ('leading_zero_col', pa.string()), ('numeric_col', pa.int64()), ('text_col', pa.string()), # Add all remaining columns with their correct types ]) table = pa.Table.from_pandas(df, schema=custom_schema) pq.write_table(table, 'output.parquet')
2. Convert the Column to Integer Type (Lose Leading Zeros)
If leading zeros are just formatting and you need the values as integers, convert the column to a numeric type first. Note: leading zeros will be lost (integers don’t store formatting):
# Convert to integers (raises error if non-numeric values exist) df['leading_zero_col'] = pd.to_numeric(df['leading_zero_col'], errors='raise') # Or handle errors gracefully (replace invalid values with NaN) # df['leading_zero_col'] = pd.to_numeric(df['leading_zero_col'], errors='coerce') # Write to Parquet without issues table = pa.Table.from_pandas(df) pq.write_table(table, 'output.parquet')
Quick Tip
Always double-check dtypes for columns with mixed-looking values (like string-formatted numbers) before converting to Parquet. Explicitly setting types avoids PyArrow’s automatic inference causing unexpected errors.
内容的提问来源于stack exchange,提问作者Carlos P Ceballos

