You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

带时区的Pandas datetime64[ns]转Arrow后Java读取为纳秒数是否正常?

带时区的Timestamp列PyArrow写入后Java读取显示为纳秒数的问题

问题场景

我有一个包含带时区datetime列的Pandas DataFrame,该列数据如下:

0   2020-01-01 23:00:57+00:00
1   2021-01-01 23:00:57+00:00
2   2022-01-01 23:00:57+00:00
3   2023-01-01 23:00:57+00:00
4   2024-01-01 23:00:57+00:00

转换为Arrow表后的结构和数据为:

pyarrow.Table
subject: string
timestamp: timestamp[ns, tz=UTC]
col1: string
----
subject: [["s1","s2","s3","s4","s5"]]
timestamp: [[2020-01-01 18:18:18.000000000,2021-01-01 18:18:18.000000000,2022-01-01 18:18:18.000000000,2023-01-01 18:18:18.000000000,2024-01-01 18:18:18.000000000]]
col1: [["71264","71264","71264","71264","71264"]]

对应的Arrow Schema:

subject: string
timestamp: timestamp[ns, tz=UTC]
col1: string
-- schema metadata --
pandas: '{"index_columns": [], "column_indexes": [], "columns": [{"name":' + 565

通过以下PyArrow代码将DataFrame写入Arrow文件:

import pyarrow as pa
writer = pa.ipc.new_file("/data.arrow", schema)
writer.write(table)

再用Java代码读取该文件时,发现timestamp列被读取为纳秒格式的长整数,而非可读的日期时间格式:

File arrowFile = new File("data.arrow");
FileInputStream fileInputStream = new FileInputStream(arrowFile);
SeekableReadChannel seekableReadChannel = new SeekableReadChannel(fileInputStream.getChannel());
ArrowFileReader arrowFileReader = new ArrowFileReader(seekableReadChannel, new RootAllocator(Integer.MAX_VALUE));

VectorSchemaRoot root  = arrowFileReader.getVectorSchemaRoot(); // get root 
Schema schema = root.getSchema(); // get schema

for (ArrowBlock arrowBlock : arrowFileReader.getRecordBlocks()) {
    arrowFileReader.loadRecordBatch(arrowBlock);
    VectorSchemaRoot schemaRoot = arrowFileReader.getVectorSchemaRoot();

    for (int i = 0; i < schemaRoot.getFieldVectors().size(); i++) {
        FieldVector fieldVector = schemaRoot.getFieldVectors().get(i);
        // 结果为TimeStampNanoTZVector[1577902698000000000, 1609525098000000000, 1641061098000000000, 1672597098000000000, 1704133098000000000] 
        // 而非日期时间格式
    }
}

进一步排查发现:当DataFrame的datetime列无时区时,Arrow表的timestamp类型为timestamp[ns],Java读取能正常得到日期时间格式;但带时区时类型为timestamp[ns, tz=UTC],Java读取结果为纳秒数。

问题解答

这不是PyArrow为提升IO效率的预期行为,而是由Arrow规范和Java Arrow库的实现特性共同决定的:

  1. Arrow的timestamp类型定义
    Arrow的timestamp类型分为两类:

    • 无时区的timestamp[ns]:存储为从epoch开始的纳秒数,Java Arrow库会自动将其转换为带默认时区的日期时间对象展示
    • 带时区的timestamp[ns, tz=UTC]:存储的本质同样是UTC时区下的epoch纳秒数,但为了明确保留时区信息,Java库会用TimeStampNanoTZVector封装,直接暴露底层数值,而非自动转换为可读格式——这是因为带时区的时间需要结合时区信息才能准确解析,库不会擅自默认转换。
  2. PyArrow的写入行为符合规范
    PyArrow将带时区的datetime列映射为Arrow的带时区timestamp类型,是完全符合Arrow数据规范的行为,和IO效率无关,核心是为了准确保留时区信息,避免跨系统读取时出现时区偏差。

  3. Java端的可读格式转换方案
    如果需要将纳秒数转换为可读的日期时间格式,可以通过TimeStampNanoTZVector的API获取时区信息,再手动转换:

    // 示例:将带时区的timestamp数值转换为UTC可读格式
    TimeStampNanoTZVector tzVector = (TimeStampNanoTZVector) fieldVector;
    // 获取时区信息
    ZoneId zoneId = ZoneId.of(tzVector.getField().getMetadata().get("timezone"));
    for (int j = 0; j < tzVector.getValueCount(); j++) {
        long nanoValue = tzVector.get(j);
        // 转换为Instant对象
        Instant instant = Instant.ofEpochSecond(nanoValue / 1_000_000_000, nanoValue % 1_000_000_000);
        // 转换为对应时区的日期时间
        ZonedDateTime zonedDateTime = ZonedDateTime.ofInstant(instant, zoneId);
        // 格式化输出
        System.out.println(zonedDateTime.format(DateTimeFormatter.ISO_ZONED_DATE_TIME));
    }
    

内容的提问来源于stack exchange,提问作者Qiang Yao

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 16:14:52