使用dplyr按周期统计Reddit中企业提及占比的问题
问题:计算俄勒冈州最低工资政策前后各企业提及占比
我有一组Reddit数据,想要统计俄勒冈州最低工资政策实施前后,不同私营企业的提及占比。数据中directed_to_whom变量记录了单条帖子提及的私企名称,treatment_announcement变量用pre/post标记政策前后阶段。目标是得到各企业在两个阶段的提及占比,用于制作对比柱状图。
数据结构示例
dput(df[1:6,c(2, 3)]) # 打印指定列的数据示例
数据样例:
structure(list(directed_to_whom = c("nike", "nike", "amazon", "walmart", "walmart", "walmart"), treatment_announcement = c("pre", "pre", "pre", "pre", "post", "post")), class = c("tbl_df", "tbl", "data.frame"), row.names = c(NA, -6L)) -> df
当前代码及问题
我原本用以下代码计算占比:
df2 <- df %>% select(directed_to_whom, treatment_announcement) %>% group_by(treatment_announcement) %>% summarise(total_posts = n(), entity_count = sum(directed_to_whom == "nike"), entity_count = sum(directed_to_whom == "amazon"), entity_count = sum(directed_to_whom == "walmart"), entity_share = entity_count/total_posts * 100)
这段代码没有报错,但结果不符合预期:它只统计了最后一家企业(沃尔玛)在政策前后的占比,而不是每个企业分别在两个阶段的占比。当前输出示例:
dput(df2[1:2,c(1,2, 3,4)])
输出结果:
structure(list(treatment_announcement = c("post", "pre"), total_posts = c(1013L, 179L), entity_count = c(152L, 26L), entity_share = c(15.004935834156, 14.5251396648045)), row.names = c(NA, -2L), class = c("tbl_df", "tbl", "data.frame"))
我期望得到的结构化结果如下:
| directed_to_whom | treatment_announcement | entity_share |
|---|---|---|
| walmart | pre | 45% |
| amazon | pre | 10% |
| nike | pre | 45% |
| walmart | post | 60% |
| amazon | post | 15% |
| nike | post | 25% |
解决方案
问题出在分组方式和变量赋值上——你只按treatment_announcement分组,并且重复赋值entity_count导致前面的计算被覆盖。正确的做法是同时按treatment_announcement和directed_to_whom分组,先统计每个企业在各阶段的提及次数,再计算占比:
方法一:分步计算
# 1. 统计每个企业在各阶段的提及次数,同时得到各阶段总帖子数 df_summary <- df %>% group_by(treatment_announcement, directed_to_whom) %>% summarise(entity_count = n(), .groups = "drop_last") %>% mutate(total_posts = sum(entity_count)) %>% ungroup() %>% # 2. 计算占比 mutate(entity_share = (entity_count / total_posts) * 100) %>% # 可选:保留需要的列 select(directed_to_whom, treatment_announcement, entity_share)
方法二:使用prop.table简化计算
df_summary <- df %>% group_by(treatment_announcement, directed_to_whom) %>% summarise(entity_count = n(), .groups = "drop") %>% group_by(treatment_announcement) %>% mutate(entity_share = prop.table(entity_count) * 100) %>% ungroup() %>% select(directed_to_whom, treatment_announcement, entity_share)
验证结果
运行上述代码后,会得到你期望的长格式数据,每个企业对应政策前后的占比,可直接用于绘制分组柱状图(比如用ggplot2):
library(ggplot2) ggplot(df_summary, aes(x = directed_to_whom, y = entity_share, fill = treatment_announcement)) + geom_col(position = "dodge") + labs(x = "企业名称", y = "提及占比(%)", fill = "政策阶段") + theme_minimal()
内容的提问来源于stack exchange,提问作者nesta1990
相关产品推荐
相关产品推荐

