如何用R对predicted_forecast_date按4天分箱计算平均温度?
用R实现按4天对预测日期分箱并计算分箱平均温度
需求说明
需要将predicted_forecast_date中的日期按4天为一组进行分箱,计算每个分箱的平均温度,并为每条数据新增对应的分箱平均温度列。
补充数据
created_forecast_date,predicted_forecast_date,daily_avg_temp 10/7/17,10/7/17,51.16868 10/7/17,10/8/17,62.60385 10/7/17,10/9/17,60.01031 10/7/17,10/10/17,59.02917 10/7/17,10/11/17,47.96719 10/7/17,10/12/17,45.26833 10/7/17,10/13/17,47.89635 10/7/17,10/14/17,55.4725 10/7/17,10/15/17,43.23625 10/7/17,10/16/17,37.19208 10/7/17,10/17/17,42.74482 10/7/17,10/18/17,40.49875 10/7/17,10/19/17,41.7275 10/7/17,10/20/17,41.88375 10/7/17,10/21/17,42.08875 10/7/17,10/22/17,43.45625 10/8/17,10/8/17,62.45715 10/8/17,10/9/17,59.4224 10/8/17,10/10/17,61.53281 10/8/17,10/11/17,48.98281 10/8/17,10/12/17,49.08937 10/8/17,10/13/17,47.71719 10/8/17,10/14/17,56.45708 10/8/17,10/15/17,42.81 10/8/17,10/16/17,44.59833 10/8/17,10/17/17,50.08292 10/8/17,10/18/17,41.28101 10/8/17,10/19/17,41.775 10/8/17,10/20/17,47.6075 10/8/17,10/21/17,50.31375 10/8/17,10/22/17,54.40625 10/8/17,10/23/17,50.16125 10/8/17,10/24/17,50.08625
用户尝试的代码片段
dput(head(DailyAverages)) structure(list(created_forecast_date = structure(c(17446, 17446, 17446, 17446, 17446, 17446), class = "Date"), predicted_forecast_date = structure(c(17446, 17447, 17448, 17449, 17450, 17451), class = "Date"), daily_avg_temp = c(51.1686805555556, 62.6038541666667, 60.0103125, 59.0291666666667, 47.9671875, 45.2683333333333 ), daily_min_temp = c(41.195, 57.015, 54.7475, 52.83, 41.1975, 36.3125), daily_max_temp = c(64.48, 67.19, 65.0025, 68.43, 54.84, 57.9875), daily_avg_dew_point = c(43.2444444444444, 59.1958333333333, 57.2585416666667, 54.3040625, 35.4146875, 34.0580208333333), daily_min_dew_point = c(38.21, 51.26, 54.2025, 46.275, 32.315, 31.895), daily_max_dew_point = c(50.9325, 64.125, 59.4, 58.595, 43.015, 38.61), daily_avg_pressure = c(1019.49972222222, 1009.33604166667, 1013.826875, 1013.21354166667, 1023.82, 1030.55802083333), daily_min_pressure = c(1015.5475, 1003.29, 1008.15, 1012.1925, 1016.535, 1028.54), daily_max_pressure = c(1021.97333333333, 1015.44, 1016.66, 1015.2775, 1028.4825, 1032.4875), daily_avg_ground_pressure = c(992.186770833333, 982.6509375, 986.999375, 986.559270833333, 996.352083333333, 1002.77677083333), daily_min_ground_pressure = c(988.4975, 976.5325, 981.6625, 985.6275, 989.6175, 1001.425), daily_max_ground_pressure = c(994.4, 988.3075, 989.88, 988.5125, 1000.9675, 1004.3725), daily_avg_humidity = c(76.0541666666667, 88.814375, 90.9704166666667, 86.1204166666667, 62.5190625, 66.6915625), daily_avg_clouds = c(20.4097222222222, 84.90625, 67.21875, 40.5416666666667, 9.78125, 10.8020833333333), daily_avg_wind_speed = c(6.20059027777778, 11.9998958333333, 5.4228125, 4.80208333333333, 8.73354166666667, 3.544375), daily_avg_rain = c(0, 11.8425, 0.5625, 0.3825, 0, 0), daily_avg_accumulated = c(0, 11.8425, 0.5625, 0.3825, 0, 0)), class = c("grouped_df", "tbl_df", "tbl", "data.frame" ), row.names = c(NA, -6L), groups = structure(list(created_forecast_date = structure(17446, class = "Date"), .rows = structure(list(1:6), ptype = integer(0), class = c("vctrs_list_of", "vctrs_vctr", "list"))), class = c("tbl_df", "tbl", "data.frame" ), row.names = c(NA, -1L), .drop = TRUE))
解决方案
可以借助dplyr包完成分组、分箱、计算合并的操作,步骤如下:
1. 加载包并处理日期格式
如果数据中的日期是字符类型,先转换为Date类型;若已为日期类型(如用户提供的DailyAverages)可跳过此步骤:
library(dplyr) # 读取补充数据并转换日期格式(示例) df <- read.csv(text = "created_forecast_date,predicted_forecast_date,daily_avg_temp 10/7/17,10/7/17,51.16868 ...") # 替换为完整补充数据 df <- df %>% mutate(across(c(created_forecast_date, predicted_forecast_date), as.Date, format = "%m/%d/%y"))
2. 分箱并计算分箱平均温度
核心逻辑是按创建日期分组,对每组内的预测日期按4天划分区间,再计算每个区间的平均温度:
result <- DailyAverages %>% ungroup() %>% # 若原数据为分组状态,先取消分组 group_by(created_forecast_date) %>% mutate( # 计算当前预测日期与组内首个预测日期的天数差 days_since_first = as.integer(predicted_forecast_date - min(predicted_forecast_date)), # 生成4天为单位的分箱标签 temp_bin = floor(days_since_first / 4) + 1 # 分箱编号从1开始 ) %>% group_by(created_forecast_date, temp_bin) %>% mutate( # 计算当前分箱的平均温度 bin_avg_temp = mean(daily_avg_temp, na.rm = TRUE) ) %>% ungroup() %>% select(-days_since_first) # 可选:移除中间变量
3. 验证结果
处理后的数据会新增temp_bin(分箱编号)和bin_avg_temp(对应分箱的平均温度)两列,符合需求格式。
内容的提问来源于stack exchange,提问作者broccolifarmer
相关产品推荐
相关产品推荐

