R长格式数据按组计算加权均值并保留长格式结构的实现方法
实现方案
你可以通过简单的长宽转换先计算分组加权均值,再把计算结果作为新行合并回原长表,全程使用tidyverse语法实现逻辑非常清晰:
library(tidyverse) # 第一步:按教育水平分组计算加权平均收入 group_weighted_mean <- dft_w %>% # 将收入、权重从行维度转成列维度,方便计算 pivot_wider(names_from = Item, values_from = Amount) %>% group_by(education) %>% summarise( Amount = weighted.mean(income, weight), Item = "weighted_income", .groups = "drop" ) # 第二步:把加权均值匹配到对应观测,合并回原数据集后按要求排序 result <- dft_w %>% distinct(obs, education) %>% left_join(group_weighted_mean, by = "education") %>% bind_rows(dft_w, .) %>% arrange(obs, factor(Item, levels = c("income", "weight", "weighted_income")))
运行print(result, n=12)即可得到你期望的输出:
# A tibble: 12 × 4 obs education Item Amount <int> <chr> <fct> <dbl> 1 1 A income 1000 2 1 A weight 10 3 1 A weighted_income 1167. 4 2 B income 2000 5 2 B weight 1 6 2 B weighted_income 1583. 7 3 B income 1500 8 3 B weight 5 9 3 B weighted_income 1583. 10 4 A income 2000 11 4 A weight 2 12 4 A weighted_income 1167.
内容的提问来源于stack exchange,提问作者Eric Fail
相关产品推荐
相关产品推荐

