按月计算分类变量子组占比:negative_count为NA问题排查
问题:按月统计情感数据时部分月份出现NA值
数据集结构
我的数据集结构如下(展示前6行):
dput(head(control_group[(1:10)]))
输出:
structure(list(post = c(date = structure(c(1299024000, 1299024000, 1299024000, 1299024000, 1299024000, 1299024000), tzone = "UTC", class = c("POSIXct", "POSIXt")),""), sentiment_human_coded = c("negative", "neutral", "negative", "neutral", "neutral", "negative"), economic_demand_complaint = c(1, 1, 1, 1, 1, 1), socio_egotropic = c("sociotropic", "sociotropic", "sociotropic", "sociotropic", "sociotropic", "sociotropic"), collective_action = c(1, 1, 1, 1, 1, 1), treatment_details = c("pre", "pre", "pre", "pre", "pre", "pre"), treatment_implementation = c("pre", "pre", "pre", "pre", "pre", "pre"), month_year = structure(c(2011.16666666667, 2011.16666666667, 2011.16666666667, 2011.16666666667, 2011.16666666667, 2011.16666666667), class = "yearmon")), row.names = c(NA, -6L), class = c("tbl_df", "tbl", "data.frame"))
统计需求
我需要按月统计以下指标:
sentiment_count:当月总条目数negative_count:当月负面情感条目数negative_share:当月负面情感占比
期望输出格式示例:
month_year sentiment_count negative_count negative_share April 2022. 300 100 33.3% May 2022. 400 100 25%
尝试的代码
最初的代码不符合需求:
graph <- control_group %>% group_by(sentiment_human_coded, month_year) %>% mutate(sentiment_month_count=n()) %>% #count of sentiment by month group_by(month_year) %>% mutate(month_year_count=n()) %>% ###total count per month mutate(sentiment_percentage = sentiment_month_count/month_year_count*100) #percentage
采用harre提供的简洁方案:
control_group %>% group_by(month_year) |> summarise(sentiment_count = n(), negative_count = sum(sentiment_human_coded == "negative"), negative_share = negative_count/sentiment_count * 100)
但运行后发现,2011年3月的negative_count和negative_share都为NA,而我确认该月实际有123条负面情感案例。
问题原因与解决办法
原因分析
出现NA的核心原因是2011年3月的sentiment_human_coded列存在缺失值(NA)。当执行sum(sentiment_human_coded == "negative")时,只要列中有NA,sentiment_human_coded == "negative"的结果就会包含NA值,而sum()函数遇到NA时会直接返回NA,导致后续的negative_share也变成NA。
验证方法
你可以先检查该月数据中sentiment_human_coded的缺失情况:
library(zoo) # 需要加载yearmon依赖包 control_group %>% filter(month_year == as.yearmon("2011-03")) %>% count(is.na(sentiment_human_coded))
解决办法
在sum()函数中添加na.rm = TRUE参数,忽略缺失值的影响,这样只会统计明确为"negative"的条目:
control_group %>% group_by(month_year) |> summarise(sentiment_count = n(), negative_count = sum(sentiment_human_coded == "negative", na.rm = TRUE), negative_share = negative_count/sentiment_count * 100)
如果需要将negative_share格式化为百分比字符串,可以使用scales包的percent()函数:
library(scales) control_group %>% group_by(month_year) |> summarise(sentiment_count = n(), negative_count = sum(sentiment_human_coded == "negative", na.rm = TRUE), negative_share = percent(negative_count/sentiment_count, accuracy = 0.1))
内容的提问来源于stack exchange,提问作者nesta1990
相关产品推荐
相关产品推荐

