如何按参与者ID、团队分组统计文本自定义情感出现次数及排错
问题解决:按参与者分组统计情感词出现次数
问题背景
现有如下结构的数据集(含重复时间观测):
df <- data.frame( participant_ID = 1:4, TeamID = c("A", "A", "B", "B"), Text1 = c( "I shouted angrily, but then felt happy. Now I am very happy.", "He is always calm and never gets angry", "She laughed, making everyone happy", "I'm feeling sad today" ), )
自定义情感词典:
emotions <- list( anger = c("angry", "shouted"), happiness = c("happy", "laughed"), sadness = c("sad") )
目标是按participant_ID和TeamID分组,统计每个情感类别的出现次数,生成如下格式的结果:
ID TeamID anger happiness sadness 1 1 A 1 2 0 2 2 A 1 0 0 3 3 B 0 2 0 4 4 B 0 0 1
原代码错误分析
原代码出现arguments imply differing number of rows: 1, 0错误,核心原因有两个:
- 列名误用:原数据框中参与者ID列是
participant_ID,但循环里错误调用了df$ID[i],导致取到空值。 - 计数逻辑缺陷:
grepl仅返回是否存在匹配(TRUE/FALSE),sum(grepl(...))只能统计是否出现,无法计算重复出现的次数(比如第一个文本里"happy"出现2次,原函数只会返回1)。
修复方案
方案1:修正原循环代码
先修正列名和计数逻辑,使用stringr::str_count统计每个关键词的出现次数总和:
# 加载所需包 library(stringr) # 初始化结果数据框 emotion_count_df <- data.frame(ID = integer(0), TeamID = character(0), Emotion = character(0), Count = integer(0)) # 修正后的情感计数函数 count_emotions <- function(text, emotions) { text <- tolower(text) # 用单词边界\b避免部分匹配(比如"angry"不会匹配"angrily"的部分) counts <- sapply(emotions, function(terms) sum(str_count(text, paste0("\\b", terms, "\\b")))) return(counts) } # 循环统计 for (i in 1:nrow(df)) { counts <- count_emotions(df$Text1[i], emotions) for (j in 1:length(emotions)) { emotion_count_df <- rbind(emotion_count_df, data.frame(ID = df$participant_ID[i], TeamID = df$TeamID[i], Emotion = names(emotions)[j], Count = counts[j])) } } # 转换为宽格式(目标格式) final_df <- reshape2::dcast(emotion_count_df, ID + TeamID ~ Emotion, value.var = "Count", fill = 0)
方案2:高效tidyverse方案(适合大数据量)
避免循环,使用dplyr和purrr处理,效率更高,尤其适合含时间维度的重复观测数据:
library(tidyverse) # 定义计数函数,返回每个情感类别的次数 count_emotions_tidy <- function(text, emotions) { text <- tolower(text) map_dfr(emotions, ~sum(str_count(text, paste0("\\b", .x, "\\b"))), .id = "Emotion") %>% rename(Count = value) } # 分组统计并转换为宽格式 final_df <- df %>% rowwise() %>% mutate(counts = list(count_emotions_tidy(Text1, emotions))) %>% unnest(counts) %>% ungroup() %>% pivot_wider(names_from = Emotion, values_from = Count, values_fill = 0) %>% rename(ID = participant_ID)
最终结果
运行上述代码后,得到目标格式的结果:
> final_df # A tibble: 4 × 5 ID TeamID anger happiness sadness <int> <chr> <int> <int> <int> 1 1 A 1 2 0 2 2 A 1 0 0 3 3 B 0 2 0 4 4 B 0 0 1
内容的提问来源于stack exchange,提问作者user22571454
相关产品推荐
相关产品推荐

