You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用dplyr按周期统计Reddit中企业提及占比的问题

问题:计算俄勒冈州最低工资政策前后各企业提及占比

我有一组Reddit数据,想要统计俄勒冈州最低工资政策实施前后,不同私营企业的提及占比。数据中directed_to_whom变量记录了单条帖子提及的私企名称,treatment_announcement变量用pre/post标记政策前后阶段。目标是得到各企业在两个阶段的提及占比,用于制作对比柱状图。

数据结构示例

dput(df[1:6,c(2, 3)]) # 打印指定列的数据示例

数据样例:

structure(list(directed_to_whom = c("nike", "nike", 
"amazon", "walmart", "walmart", "walmart"), treatment_announcement = c("pre", 
"pre", "pre", "pre", "post", "post")), class = c("tbl_df", "tbl", 
"data.frame"), row.names = c(NA, -6L)) -> df

当前代码及问题

我原本用以下代码计算占比:

df2  <- df %>%
  select(directed_to_whom, treatment_announcement) %>%
  group_by(treatment_announcement) %>%
  summarise(total_posts = n(),
            entity_count = sum(directed_to_whom == "nike"),
            entity_count = sum(directed_to_whom == "amazon"),
            entity_count = sum(directed_to_whom == "walmart"),
            entity_share = entity_count/total_posts * 100) 

这段代码没有报错,但结果不符合预期:它只统计了最后一家企业(沃尔玛)在政策前后的占比,而不是每个企业分别在两个阶段的占比。当前输出示例:

dput(df2[1:2,c(1,2, 3,4)])

输出结果:

structure(list(treatment_announcement = c("post", "pre"), total_posts = c(1013L, 
179L), entity_count = c(152L, 26L), entity_share = c(15.004935834156, 
14.5251396648045)), row.names = c(NA, -2L), class = c("tbl_df", 
"tbl", "data.frame"))

我期望得到的结构化结果如下:

directed_to_whomtreatment_announcemententity_share
walmartpre45%
amazonpre10%
nikepre45%
walmartpost60%
amazonpost15%
nikepost25%

解决方案

问题出在分组方式和变量赋值上——你只按treatment_announcement分组,并且重复赋值entity_count导致前面的计算被覆盖。正确的做法是同时按treatment_announcement和directed_to_whom分组,先统计每个企业在各阶段的提及次数,再计算占比:

方法一:分步计算

# 1. 统计每个企业在各阶段的提及次数,同时得到各阶段总帖子数
df_summary <- df %>%
  group_by(treatment_announcement, directed_to_whom) %>%
  summarise(entity_count = n(), .groups = "drop_last") %>%
  mutate(total_posts = sum(entity_count)) %>%
  ungroup() %>%
  # 2. 计算占比
  mutate(entity_share = (entity_count / total_posts) * 100) %>%
  # 可选:保留需要的列
  select(directed_to_whom, treatment_announcement, entity_share)

方法二:使用prop.table简化计算

df_summary <- df %>%
  group_by(treatment_announcement, directed_to_whom) %>%
  summarise(entity_count = n(), .groups = "drop") %>%
  group_by(treatment_announcement) %>%
  mutate(entity_share = prop.table(entity_count) * 100) %>%
  ungroup() %>%
  select(directed_to_whom, treatment_announcement, entity_share)

验证结果

运行上述代码后,会得到你期望的长格式数据,每个企业对应政策前后的占比,可直接用于绘制分组柱状图(比如用ggplot2):

library(ggplot2)
ggplot(df_summary, aes(x = directed_to_whom, y = entity_share, fill = treatment_announcement)) +
  geom_col(position = "dodge") +
  labs(x = "企业名称", y = "提及占比(%)", fill = "政策阶段") +
  theme_minimal()

内容的提问来源于stack exchange,提问作者nesta1990

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 14:12:47