You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按参与者ID、团队分组统计文本自定义情感出现次数及排错

问题解决:按参与者分组统计情感词出现次数

问题背景

现有如下结构的数据集(含重复时间观测):

df <- data.frame(
  participant_ID = 1:4, 
  TeamID = c("A", "A", "B", "B"),
  Text1 = c(
      "I shouted angrily, but then felt happy. Now I am very happy.",
      "He is always calm and never gets angry",
      "She laughed, making everyone happy",
      "I'm feeling sad today"
    ), 
)

自定义情感词典:

emotions <- list(
  anger = c("angry", "shouted"),
  happiness = c("happy", "laughed"),
  sadness = c("sad")
)

目标是按participant_ID和TeamID分组,统计每个情感类别的出现次数,生成如下格式的结果:

ID TeamID anger happiness sadness 
1  1      A     1         2       0
2  2      A     1         0       0
3  3      B     0         2       0
4  4      B     0         0       1

原代码错误分析

原代码出现arguments imply differing number of rows: 1, 0错误,核心原因有两个:

  • 列名误用:原数据框中参与者ID列是participant_ID,但循环里错误调用了df$ID[i],导致取到空值。
  • 计数逻辑缺陷:grepl仅返回是否存在匹配(TRUE/FALSE),sum(grepl(...))只能统计是否出现,无法计算重复出现的次数(比如第一个文本里"happy"出现2次,原函数只会返回1)。

修复方案

方案1:修正原循环代码

先修正列名和计数逻辑,使用stringr::str_count统计每个关键词的出现次数总和:

# 加载所需包
library(stringr)

# 初始化结果数据框
emotion_count_df <- data.frame(ID = integer(0), TeamID = character(0), Emotion = character(0), Count = integer(0))

# 修正后的情感计数函数
count_emotions <- function(text, emotions) {
  text <- tolower(text)
  # 用单词边界\b避免部分匹配(比如"angry"不会匹配"angrily"的部分)
  counts <- sapply(emotions, function(terms) sum(str_count(text, paste0("\\b", terms, "\\b"))))
  return(counts)
}

# 循环统计
for (i in 1:nrow(df)) {
  counts <- count_emotions(df$Text1[i], emotions)
  for (j in 1:length(emotions)) {
    emotion_count_df <- rbind(emotion_count_df, 
                              data.frame(ID = df$participant_ID[i], 
                                         TeamID = df$TeamID[i],
                                         Emotion = names(emotions)[j], 
                                         Count = counts[j]))
  }
}

# 转换为宽格式(目标格式)
final_df <- reshape2::dcast(emotion_count_df, ID + TeamID ~ Emotion, value.var = "Count", fill = 0)

方案2:高效tidyverse方案(适合大数据量)

避免循环,使用dplyr和purrr处理,效率更高,尤其适合含时间维度的重复观测数据:

library(tidyverse)

# 定义计数函数,返回每个情感类别的次数
count_emotions_tidy <- function(text, emotions) {
  text <- tolower(text)
  map_dfr(emotions, ~sum(str_count(text, paste0("\\b", .x, "\\b"))), .id = "Emotion") %>%
    rename(Count = value)
}

# 分组统计并转换为宽格式
final_df <- df %>%
  rowwise() %>%
  mutate(counts = list(count_emotions_tidy(Text1, emotions))) %>%
  unnest(counts) %>%
  ungroup() %>%
  pivot_wider(names_from = Emotion, values_from = Count, values_fill = 0) %>%
  rename(ID = participant_ID)

最终结果

运行上述代码后,得到目标格式的结果:

> final_df
# A tibble: 4 × 5
     ID TeamID anger happiness sadness
  <int> <chr>  <int>     <int>   <int>
1     1 A          1         2       0
2     2 A          1         0       0
3     3 B          0         2       0
4     4 B          0         0       1

内容的提问来源于stack exchange,提问作者user22571454

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 02:58:18