如何按含重复值的向量索引DataFrame以计算加权平均值?
解决重复年月向量索引DataFrame的问题
核心问题
用%in%索引会自动去重,仅返回唯一匹配的数值,无法保留采样的重复次数,因此需要用能保留顺序和重复项的索引方法。
具体解决方案
1. 用match()函数实现精准匹配
match()会按目标向量的顺序,逐个定位每个元素在DataFrame年月列中的位置,返回对应数值,完美保留重复采样的次数。
示例代码:
# 构造示例月度数据 monthly_data <- data.frame( year_month = c("2000-01", "2000-02", "2000-03"), value = c(10, 20, 30) ) # 包含重复年月的采样向量 sample_dates <- c("2000-01", "2000-01", "2000-01", "2000-01", "2000-02") # 错误方法:%in%仅返回唯一值 wrong_values <- monthly_data$value[monthly_data$year_month %in% sample_dates] # 输出结果:10 20(仅两个唯一值) # 正确方法:用match()保留重复项 correct_values <- monthly_data$value[match(sample_dates, monthly_data$year_month)] # 输出结果:10 10 10 10 20(完全对应采样的重复次数)
2. 基于因子的索引(可选)
如果year_month列是因子类型,且采样向量的元素均为因子的水平,直接用采样向量索引即可:
# 将year_month转为因子 monthly_data$year_month <- as.factor(monthly_data$year_month) # 直接用采样向量索引 correct_values <- monthly_data$value[sample_dates]
计算加权平均值
得到重复数值向量后,直接用mean()就能得到按采样次数加权的平均值(每个重复值等价于权重1):
weighted_avg <- mean(correct_values) # 计算结果:(10*4 + 20*1)/5 = 12
如果需要手动基于采样次数占比计算权重,也可以这么写:
# 统计每个年月的采样次数占比 date_weights <- table(sample_dates) / length(sample_dates) # 匹配权重并计算加权和 weighted_avg <- sum(monthly_data$value * date_weights[monthly_data$year_month])
内容的提问来源于stack exchange,提问作者Jake L
相关产品推荐
相关产品推荐

