You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Pandas DataFrame中YYYYMMDD格式的字符串列转换为Parquet的DATE类型

Solution: Map YYYYMMDD String to Parquet DATE Type

Alright, let's fix this. The issue with your current approach is that converting the datetime back to a string tells Parquet to treat it as a STRING type—exactly what you're seeing. To get the DATE type in your Parquet schema (showing as optional int32 exam_date (DATE);), you need to map your date data to Parquet's native DATE type, which stores dates as an int32 representing days since 1970-01-01.

Here's a step-by-step implementation using PyArrow (it handles the type conversion to Parquet's DATE correctly):

Step 1: Import Required Libraries

import pandas as pd
import pyarrow as pa
import pyarrow.parquet as pq

Step 2: Convert the String Column to Date Objects

First, turn your YYYYMMDD string column into Python date objects (this gives us a type PyArrow can easily map to Parquet's DATE):

# Example DataFrame (replace with your actual data)
final_calc_df = pd.DataFrame({'exam_date': ['20201130', '20210101', '20221231']})

# Convert string to datetime, then extract the pure date component
final_calc_df['exam_date'] = pd.to_datetime(final_calc_df['exam_date'], format='%Y%m%d').dt.date

Step 3: Define the Parquet Schema

Explicitly specify that exam_date should be a date32 type (which maps directly to Parquet's DATE type):

# Build the schema (add other columns here if your DataFrame has more fields)
schema = pa.schema([
    ('exam_date', pa.date32())
])

Step 4: Save as Parquet with the Correct Schema

Convert your DataFrame to a PyArrow Table (applying the schema) and write it to Parquet:

# Convert DataFrame to PyArrow Table to enforce the schema
table = pa.Table.from_pandas(final_calc_df, schema=schema)

# Write to Parquet file
pq.write_table(table, 'myfile.parquet')

Alternatively, you can use Pandas' to_parquet method directly with the PyArrow engine:

final_calc_df.to_parquet(
    'myfile.parquet',
    engine='pyarrow',
    schema=schema
)

Verify the Schema

Run your Parquet tools command to confirm:

java -jar parquet-tools.jar schema myfile.parquet

You'll see the desired output:

message schema {
  optional int32 exam_date (DATE);
}

Why This Works

  • pa.date32() is PyArrow's implementation of Parquet's standard DATE type, which stores dates as an int32 (count of days since 1970-01-01).
  • By avoiding converting back to a string and instead using date objects with an explicit schema, we tell Parquet to use the correct DATE type instead of STRING.

内容的提问来源于stack exchange,提问作者Abhishek Patil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.01 03:17:32