You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PyArrow导入含微秒时间戳的CSV至Parquet时%f解析失败求助

问题:PyArrow解析CSV含微秒的时间戳失败,需保留数据完整性

我尝试用PyArrow将CSV数据导入Parquet,通过ConvertOptions指定列类型和时间戳解析规则,但发现PyArrow无法识别用于解析微秒的%f格式符。移除数据中的微秒并删除timestamp_parsers里的%f后代码可正常运行,但我需要保留数据完整性,因此该方案不可行。想确认这是PyArrow的Bug还是操作疏漏,寻求解决方案。

CSV示例

time,data
01-11-19 10:11:56.132,xxx

代码示例

import pyarrow as pa
from pyarrow import csv
from pyarrow import parquet


convert_dict = {
    'time': pa.timestamp('us', None),
    'data': pa.string()
}

convert_options = csv.ConvertOptions(
    column_types=convert_dict
    , strings_can_be_null=True
    , quoted_strings_can_be_null=True
    , timestamp_parsers=['%d-%m-%y %H:%M:%S.%f']
)

table = csv.read_csv('test.csv', convert_options=convert_options)
print(table)
parquet.write_table(table, 'test.parquet')

解答

这不是PyArrow的Bug,是你误用了时间格式符:PyArrow的timestamp_parsers遵循的是ICU DateFormat规范,而非Python原生的strptime规则,因此Python里表示微秒的%f在ICU规范里不具备相同含义,导致解析失败。

解决方案

将timestamp_parsers中的格式符替换为ICU规范对应的写法即可:

  • 如果你明确时间戳的小数秒是3位(毫秒),使用%d-%m-%y %H:%M:%S.%3f
  • 如果小数秒位数不固定(比如可能是3位或6位),直接使用%d-%m-%y %H:%M:%S.%f(ICU中的%f表示秒的小数部分,支持任意位数)

修改后的convert_options代码:

convert_options = csv.ConvertOptions(
    column_types=convert_dict
    , strings_can_be_null=True
    , quoted_strings_can_be_null=True
    , timestamp_parsers=['%d-%m-%y %H:%M:%S.%3f']
)

运行修改后的代码,即可正确解析带小数秒的时间戳,同时完整保留原始数据的时间精度。


内容的提问来源于stack exchange,提问作者not_a_comp_scientist

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 22:30:48