You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

按月计算分类变量子组占比:negative_count为NA问题排查

问题:按月统计情感数据时部分月份出现NA值

数据集结构

我的数据集结构如下(展示前6行):

dput(head(control_group[(1:10)]))

输出:

structure(list(post = c(date = structure(c(1299024000, 1299024000, 1299024000, 1299024000, 
1299024000, 1299024000), tzone = "UTC", class = c("POSIXct", 
"POSIXt")),""), sentiment_human_coded = c("negative", 
"neutral", "negative", "neutral", "neutral", "negative"), economic_demand_complaint = c(1, 
1, 1, 1, 1, 1), socio_egotropic = c("sociotropic", "sociotropic", 
"sociotropic", "sociotropic", "sociotropic", "sociotropic"), 
    collective_action = c(1, 1, 1, 1, 1, 1), treatment_details = c("pre", 
"pre", "pre", "pre", "pre", "pre"), treatment_implementation = c("pre", 
"pre", "pre", "pre", "pre", "pre"), month_year = structure(c(2011.16666666667, 
2011.16666666667, 2011.16666666667, 2011.16666666667, 2011.16666666667, 
2011.16666666667), class = "yearmon")), row.names = c(NA, 
-6L), class = c("tbl_df", "tbl", "data.frame"))

统计需求

我需要按月统计以下指标:

  • sentiment_count:当月总条目数
  • negative_count:当月负面情感条目数
  • negative_share:当月负面情感占比

期望输出格式示例:

month_year    sentiment_count  negative_count   negative_share
April 2022.   300               100              33.3%
May 2022.     400               100              25%

尝试的代码

最初的代码不符合需求:

graph <- control_group %>%
  group_by(sentiment_human_coded, month_year) %>%   
  mutate(sentiment_month_count=n()) %>% #count of sentiment by month
  group_by(month_year) %>% 
  mutate(month_year_count=n())  %>% ###total count per month
  mutate(sentiment_percentage = sentiment_month_count/month_year_count*100) #percentage

采用harre提供的简洁方案:

control_group %>%
  group_by(month_year) |>
  summarise(sentiment_count = n(),
            negative_count = sum(sentiment_human_coded == "negative"),
            negative_share = negative_count/sentiment_count * 100) 

但运行后发现,2011年3月的negative_count和negative_share都为NA,而我确认该月实际有123条负面情感案例。


问题原因与解决办法

原因分析

出现NA的核心原因是2011年3月的sentiment_human_coded列存在缺失值(NA)。当执行sum(sentiment_human_coded == "negative")时,只要列中有NA,sentiment_human_coded == "negative"的结果就会包含NA值,而sum()函数遇到NA时会直接返回NA,导致后续的negative_share也变成NA。

验证方法

你可以先检查该月数据中sentiment_human_coded的缺失情况:

library(zoo) # 需要加载yearmon依赖包
control_group %>% 
  filter(month_year == as.yearmon("2011-03")) %>% 
  count(is.na(sentiment_human_coded))

解决办法

在sum()函数中添加na.rm = TRUE参数,忽略缺失值的影响,这样只会统计明确为"negative"的条目:

control_group %>%
  group_by(month_year) |>
  summarise(sentiment_count = n(),
            negative_count = sum(sentiment_human_coded == "negative", na.rm = TRUE),
            negative_share = negative_count/sentiment_count * 100) 

如果需要将negative_share格式化为百分比字符串,可以使用scales包的percent()函数:

library(scales)
control_group %>%
  group_by(month_year) |>
  summarise(sentiment_count = n(),
            negative_count = sum(sentiment_human_coded == "negative", na.rm = TRUE),
            negative_share = percent(negative_count/sentiment_count, accuracy = 0.1)) 

内容的提问来源于stack exchange,提问作者nesta1990

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 09:50:42