如何在R中基于气象列按年份统计各指标组合的暴露天数
问题:按年份统计气象要素的暴露时长(出现天数)
我有按日间隔记录的气象数据,需要基于现有列mean_rh(平均相对湿度)、mean_temp(平均气温)和mean_ws(平均风速),按年份计算每个要素唯一值/要素组合的暴露时长(即该值/组合出现的天数)。例如查询2006年mean_ws为0.06的记录天数,从原始数据可知仅出现1天。
可复现示例数据
library(dplyr) # 创建示例数据框 df <- data.frame( year = c(2006, 2006, 2006, 2006, 2006, 2006, 2006, 2006, 2006, 2006, 2006, 2006, 2006, 2006, 2006, 2006, 2006, 2006, 2006, 2006), date = c("11/9/06", "12/9/06", "13/9/06", "14/9/06", "15/9/06", "16/9/06", "17/9/06", "18/9/06", "19/9/06", "20/9/06", "21/9/06", "22/9/06", "23/9/06", "24/9/06", "25/9/06", "26/9/06", "27/9/06", "28/9/06", "29/9/06", "30/9/06"), mean_rh = c(69.49, 69.05, 68.47, 65.99, 65.53, 67.07, 72.42, 91.79, 90.13, 76.88, 67.68, 61.11, 70.19, 66.88, 76.63, 89.09, 84.97, 80.08, 79.82, 82.81), mean_temp = c(18.05, 20.23, 20.93, 20.58, 18.64, 18.74, 19.13, 15.98, 14.62, 14.06, 17.03, 19.76, 18.64, 18.27, 16.96, 16.31, 14.97, 15.69, 16.34, 16.02), mean_ws = c(0.84, 0.33, 0.46, 0.79, 2.25, 1.95, 0.13, 0.06, 1.10, 0.90, 1.10, 1.20, 0.33, 0.36, 0.27, 0.66, 0.22, 0.62, 0.33, 0.18) ) # 将日期列转换为日期格式 df$date <- as.Date(df$date, format = "%d/%m/%y")
当前错误代码及问题
我尝试用以下代码统计温度的暴露时长,但计算结果与原始数据不符,需要修正代码,同时支持按mean_ws、mean_rh单独分组或三者组合分组统计:
# 计算每个mean_temp的暴露时长(错误代码) df2 <- df %>% group_by(year, mean_temp) %>% summarise(exposure_duration = difftime(max(date), min(date), units = "days")) %>% ungroup() # 转换为宽格式 df3 <- df2 %>% spread(key = mean_temp, value = exposure_duration, fill = 0)
修正方案及代码
错误原因
原代码中difftime(max(date), min(date))计算的是该要素值首次与末次出现的日期间隔,而非实际出现的天数。例如示例中mean_temp=18.64出现在15日和23日,日期间隔为8天,但实际出现天数是2天,这就是结果不符的核心原因。
正确的做法是用n()统计分组内的行数——由于数据按日记录,每行对应一天,行数即为该要素值/组合的出现天数。
修正后的代码
1. 单独统计某一要素的暴露时长(以mean_temp为例)
library(dplyr) library(tidyr) # 推荐使用pivot_wider替代spread,更灵活 # 按年份+mean_temp分组,统计出现天数 df_temp_exposure <- df %>% group_by(year, mean_temp) %>% summarise(exposure_duration = n(), .groups = "drop") # .groups="drop"替代ungroup()更简洁 # 转换为宽格式(每个温度值作为列) df_temp_wide <- df_temp_exposure %>% pivot_wider(names_from = mean_temp, values_from = exposure_duration, values_fill = 0)
2. 统计mean_ws的暴露时长
只需替换分组字段:
df_ws_exposure <- df %>% group_by(year, mean_ws) %>% summarise(exposure_duration = n(), .groups = "drop") df_ws_wide <- df_ws_exposure %>% pivot_wider(names_from = mean_ws, values_from = exposure_duration, values_fill = 0)
3. 统计mean_rh的暴露时长
同理:
df_rh_exposure <- df %>% group_by(year, mean_rh) %>% summarise(exposure_duration = n(), .groups = "drop") df_rh_wide <- df_rh_exposure %>% pivot_wider(names_from = mean_rh, values_from = exposure_duration, values_fill = 0)
4. 统计三个要素组合的暴露时长
同时分组三个字段即可:
df_combined_exposure <- df %>% group_by(year, mean_rh, mean_temp, mean_ws) %>% summarise(exposure_duration = n(), .groups = "drop")
内容的提问来源于stack exchange,提问作者Ahsk
相关产品推荐
相关产品推荐

