如何在R中正确读取带时区偏移的字符型datetime(夏令时场景)
问题描述
我有一个数据集,其中包含字符型datetime列,时区为CEST/CET(欧洲中部时间),末尾标注+01:00/+02:00的时区偏移。我希望将其转换为POSIXct格式以便后续转为UTC,但在10月夏令时切换日,2点到3点的额外小时读取错误,时区偏移似乎被忽略。
源数据示例:
| datetime_CEST_CET_character |
|---|
| 2023-10-29 00:00+02:00 |
| 2023-10-29 01:00+02:00 |
| 2023-10-29 02:00+02:00 |
| 2023-10-29 02:00+01:00 |
| 2023-10-29 03:00+01:00 |
| 2023-10-29 04:00+01:00 |
我的目标是让reprex$datetime_CEST_CET_converted[4]返回"2023-10-29 02:00:00 CET"而非"2023-10-29 02:00:00 CEST"。
当前使用的代码
library(dplyr) library(lubridate) source <- data.frame( datetime_CEST_CET_character = c("2023-10-29 00:00+02:00", "2023-10-29 01:00+02:00", "2023-10-29 02:00+02:00", "2023-10-29 02:00+01:00", "2023-10-29 03:00+01:00", "2023-10-29 04:00+01:00") ) reprex <- source %>% mutate(datetime_CEST_CET_converted = as.POSIXct(datetime_CEST_CET_character, tz = "Europe/Paris"), datetime_UTC = with_tz(datetime_CEST_CET_converted, tzone = "UTC")) reprex$datetime_CEST_CET_converted[3] reprex$datetime_CEST_CET_converted[4] reprex$datetime_CEST_CET_converted[5] - hours(1)
尝试过的无效方法
我尝试移除时区偏移中的冒号,在as.POSIXct()中添加format="%Y-%m-%d %H:%M+%z",但结果返回NA:
修改后的数据源:
| datetime_CEST_CET_character |
|---|
| 2023-10-29 00:00+0200 |
| 2023-10-29 01:00+0200 |
| 2023-10-29 02:00+0200 |
| 2023-10-29 02:00+0100 |
| 2023-10-29 03:00+0100 |
| 2023-10-29 04:00+0100 |
对应的代码:
source_without_colon_in_timezone <- data.frame( datetime_CEST_CET_character = c("2023-10-29 00:00+0200", "2023-10-29 01:00+0200", "2023-10-29 02:00+0200", "2023-10-29 02:00+0100", "2023-10-29 03:00+0100", "2023-10-29 04:00+0100") ) reprex_without_colon_in_timezone <- source_without_colon_in_timezone %>% mutate(datetime_CEST_CET_converted = as.POSIXct(datetime_CEST_CET_character, format="%Y-%m-%d %H:%M+%z", tz = "Europe/Paris"), datetime_UTC = with_tz(datetime_CEST_CET_converted, tzone = "UTC")) reprex_without_colon_in_timezone$datetime_CEST_CET_converted[3] reprex_without_colon_in_timezone$datetime_CEST_CET_converted[4] reprex_without_colon_in_timezone$datetime_CEST_CET_converted[5] - hours(1)
解决方案
问题核心是as.POSIXct指定tz参数时,会优先用该时区解析时间字符串,忽略自带的偏移量,导致夏令时切换的重复时间被错误归类。以下是两种可靠的处理方式:
方法1:用lubridate的parse_date_time解析时区偏移
先按字符串自带的偏移解析为UTC时间,再转换到目标时区:
library(dplyr) library(lubridate) source <- data.frame( datetime_CEST_CET_character = c("2023-10-29 00:00+02:00", "2023-10-29 01:00+02:00", "2023-10-29 02:00+02:00", "2023-10-29 02:00+01:00", "2023-10-29 03:00+01:00", "2023-10-29 04:00+01:00") ) reprex <- source %>% mutate( datetime_CEST_CET_converted = parse_date_time(datetime_CEST_CET_character, orders = "Ymd HMS z") %>% with_tz(tzone = "Europe/Paris"), datetime_UTC = with_tz(datetime_CEST_CET_converted, tzone = "UTC") ) # 验证结果 reprex$datetime_CEST_CET_converted[4] # 返回:"2023-10-29 02:00:00 CET"
方法2:先解析为UTC,再转换目标时区
不指定tz参数,让as.POSIXct自动识别字符串里的偏移得到UTC时间,再转换到欧洲中部时区:
reprex <- source %>% mutate( datetime_UTC = as.POSIXct(datetime_CEST_CET_character), datetime_CEST_CET_converted = with_tz(datetime_UTC, tzone = "Europe/Paris") ) reprex$datetime_CEST_CET_converted[4] # 返回:"2023-10-29 02:00:00 CET"
无效方法的原因
- 直接指定
tz="Europe/Paris"时,as.POSIXct会忽略字符串自带的偏移,按该时区解析时间,导致重复的02:00都被识别为CEST。 - 你尝试的
format="%Y-%m-%d %H:%M+%z"格式错误,正确格式应为"%Y-%m-%d %H:%M%z"(去掉前面的+),但即便修正格式,夏令时切换的重复时间仍可能被错误解析,因此优先推荐前两种方法。
内容的提问来源于stack exchange,提问作者Vincent
相关产品推荐
相关产品推荐

