You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyArrow:如何转换带AM/PM的时间列并保存至原表?

PyArrow处理带AM/PM的时间戳转换

一、读取CSV时直接完成转换(推荐)

在pyarrow.csv.read_csv中通过配置timestamp_parsers参数,指定带AM/PM的时间格式,就能直接将字符串解析为标准timestamp类型,无需后续二次转换。

注意:时间格式中的小时部分需用%I(适配12小时制),而非%H(24小时制),否则AM/PM标记无法被正确解析。

示例代码:

import pyarrow as pa
import pyarrow.csv as csv

table = csv.read_csv('my_data.csv',
                     convert_options=csv.ConvertOptions(
                         # 指定目标列类型为秒级timestamp
                         column_types={'Datetime': pa.timestamp('s')},
                         # 设置时间字符串的解析格式
                         timestamp_parsers=['%Y/%m/%d %I:%M:%S %p']
                     )
                    )

二、转换现有表中的列并替换回原表

如果已读取原始表,可通过pyarrow.compute.strptime转换目标列后,用set_column方法替换原表中的对应列:

示例代码:

import pyarrow.compute as pc

# 转换Datetime列:注意格式用%I适配12小时制
new_datetime_col = pc.strptime(
    table.column("Datetime"), 
    format='%Y/%m/%d %I:%M:%S %p', 
    unit='s'
).cast(pa.timestamp('s'))

# 获取原列的索引位置
col_index = table.schema.get_field_index("Datetime")

# 替换原列生成新表
updated_table = table.set_column(col_index, "Datetime", new_datetime_col)

三、写入Parquet文件

无论用哪种方式得到含标准timestamp的表,都可以直接写入Parquet:

import pyarrow.parquet as pq

pq.write_table(updated_table, 'output.parquet')

内容的提问来源于stack exchange,提问作者Illia Kaltovich

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 21:10:52