带时区的Pandas datetime64[ns]转Arrow后Java读取为纳秒数是否正常?
带时区的Timestamp列PyArrow写入后Java读取显示为纳秒数的问题
问题场景
我有一个包含带时区datetime列的Pandas DataFrame,该列数据如下:
0 2020-01-01 23:00:57+00:00 1 2021-01-01 23:00:57+00:00 2 2022-01-01 23:00:57+00:00 3 2023-01-01 23:00:57+00:00 4 2024-01-01 23:00:57+00:00
转换为Arrow表后的结构和数据为:
pyarrow.Table subject: string timestamp: timestamp[ns, tz=UTC] col1: string ---- subject: [["s1","s2","s3","s4","s5"]] timestamp: [[2020-01-01 18:18:18.000000000,2021-01-01 18:18:18.000000000,2022-01-01 18:18:18.000000000,2023-01-01 18:18:18.000000000,2024-01-01 18:18:18.000000000]] col1: [["71264","71264","71264","71264","71264"]]
对应的Arrow Schema:
subject: string timestamp: timestamp[ns, tz=UTC] col1: string -- schema metadata -- pandas: '{"index_columns": [], "column_indexes": [], "columns": [{"name":' + 565
通过以下PyArrow代码将DataFrame写入Arrow文件:
import pyarrow as pa writer = pa.ipc.new_file("/data.arrow", schema) writer.write(table)
再用Java代码读取该文件时,发现timestamp列被读取为纳秒格式的长整数,而非可读的日期时间格式:
File arrowFile = new File("data.arrow"); FileInputStream fileInputStream = new FileInputStream(arrowFile); SeekableReadChannel seekableReadChannel = new SeekableReadChannel(fileInputStream.getChannel()); ArrowFileReader arrowFileReader = new ArrowFileReader(seekableReadChannel, new RootAllocator(Integer.MAX_VALUE)); VectorSchemaRoot root = arrowFileReader.getVectorSchemaRoot(); // get root Schema schema = root.getSchema(); // get schema for (ArrowBlock arrowBlock : arrowFileReader.getRecordBlocks()) { arrowFileReader.loadRecordBatch(arrowBlock); VectorSchemaRoot schemaRoot = arrowFileReader.getVectorSchemaRoot(); for (int i = 0; i < schemaRoot.getFieldVectors().size(); i++) { FieldVector fieldVector = schemaRoot.getFieldVectors().get(i); // 结果为TimeStampNanoTZVector[1577902698000000000, 1609525098000000000, 1641061098000000000, 1672597098000000000, 1704133098000000000] // 而非日期时间格式 } }
进一步排查发现:当DataFrame的datetime列无时区时,Arrow表的timestamp类型为timestamp[ns],Java读取能正常得到日期时间格式;但带时区时类型为timestamp[ns, tz=UTC],Java读取结果为纳秒数。
问题解答
这不是PyArrow为提升IO效率的预期行为,而是由Arrow规范和Java Arrow库的实现特性共同决定的:
Arrow的timestamp类型定义
Arrow的timestamp类型分为两类:- 无时区的
timestamp[ns]:存储为从epoch开始的纳秒数,Java Arrow库会自动将其转换为带默认时区的日期时间对象展示 - 带时区的
timestamp[ns, tz=UTC]:存储的本质同样是UTC时区下的epoch纳秒数,但为了明确保留时区信息,Java库会用TimeStampNanoTZVector封装,直接暴露底层数值,而非自动转换为可读格式——这是因为带时区的时间需要结合时区信息才能准确解析,库不会擅自默认转换。
- 无时区的
PyArrow的写入行为符合规范
PyArrow将带时区的datetime列映射为Arrow的带时区timestamp类型,是完全符合Arrow数据规范的行为,和IO效率无关,核心是为了准确保留时区信息,避免跨系统读取时出现时区偏差。Java端的可读格式转换方案
如果需要将纳秒数转换为可读的日期时间格式,可以通过TimeStampNanoTZVector的API获取时区信息,再手动转换:// 示例:将带时区的timestamp数值转换为UTC可读格式 TimeStampNanoTZVector tzVector = (TimeStampNanoTZVector) fieldVector; // 获取时区信息 ZoneId zoneId = ZoneId.of(tzVector.getField().getMetadata().get("timezone")); for (int j = 0; j < tzVector.getValueCount(); j++) { long nanoValue = tzVector.get(j); // 转换为Instant对象 Instant instant = Instant.ofEpochSecond(nanoValue / 1_000_000_000, nanoValue % 1_000_000_000); // 转换为对应时区的日期时间 ZonedDateTime zonedDateTime = ZonedDateTime.ofInstant(instant, zoneId); // 格式化输出 System.out.println(zonedDateTime.format(DateTimeFormatter.ISO_ZONED_DATE_TIME)); }
内容的提问来源于stack exchange,提问作者Qiang Yao
相关产品推荐
相关产品推荐

