如何将Pandas DataFrame中YYYYMMDD格式的字符串列转换为Parquet的DATE类型
Alright, let's fix this. The issue with your current approach is that converting the datetime back to a string tells Parquet to treat it as a STRING type—exactly what you're seeing. To get the DATE type in your Parquet schema (showing as optional int32 exam_date (DATE);), you need to map your date data to Parquet's native DATE type, which stores dates as an int32 representing days since 1970-01-01.
Here's a step-by-step implementation using PyArrow (it handles the type conversion to Parquet's DATE correctly):
Step 1: Import Required Libraries
import pandas as pd import pyarrow as pa import pyarrow.parquet as pq
Step 2: Convert the String Column to Date Objects
First, turn your YYYYMMDD string column into Python date objects (this gives us a type PyArrow can easily map to Parquet's DATE):
# Example DataFrame (replace with your actual data) final_calc_df = pd.DataFrame({'exam_date': ['20201130', '20210101', '20221231']}) # Convert string to datetime, then extract the pure date component final_calc_df['exam_date'] = pd.to_datetime(final_calc_df['exam_date'], format='%Y%m%d').dt.date
Step 3: Define the Parquet Schema
Explicitly specify that exam_date should be a date32 type (which maps directly to Parquet's DATE type):
# Build the schema (add other columns here if your DataFrame has more fields) schema = pa.schema([ ('exam_date', pa.date32()) ])
Step 4: Save as Parquet with the Correct Schema
Convert your DataFrame to a PyArrow Table (applying the schema) and write it to Parquet:
# Convert DataFrame to PyArrow Table to enforce the schema table = pa.Table.from_pandas(final_calc_df, schema=schema) # Write to Parquet file pq.write_table(table, 'myfile.parquet')
Alternatively, you can use Pandas' to_parquet method directly with the PyArrow engine:
final_calc_df.to_parquet( 'myfile.parquet', engine='pyarrow', schema=schema )
Verify the Schema
Run your Parquet tools command to confirm:
java -jar parquet-tools.jar schema myfile.parquet
You'll see the desired output:
message schema { optional int32 exam_date (DATE); }
Why This Works
pa.date32()is PyArrow's implementation of Parquet's standard DATE type, which stores dates as an int32 (count of days since 1970-01-01).- By avoiding converting back to a string and instead using date objects with an explicit schema, we tell Parquet to use the correct DATE type instead of STRING.
内容的提问来源于stack exchange,提问作者Abhishek Patil

