如何在R语言中为数据集添加关联原数据行的情感分析列
Got it, let's solve this problem step by step—getting per-row sentiment scores tied back to your original dataset is totally doable with a couple of solid R packages. Here's how to pull it off, tailored to your data structure (case IDs + text):
First, let's assume your imported data looks like this (replace with your actual BI-loaded data):
# 加载必备工具包 library(tidyverse) # 模拟你的数据结构(实际用read_csv/read.table导入BI数据) your_data <- tibble( case_id = c("12345", "67890", "13579"), text = c("I am so happy with the service today!", "This product is terrible, I'm really disappointed.", "It's okay, not great but not bad either.") )
方法1:使用tidytext(灵活支持多种情感词典)
tidytext lets you use well-known sentiment dictionaries (like AFINN for numeric scores, or BING for positive/negative categories) and easily tie results back to your original rows.
选项A:获取数值型情感分数(AFINN词典)
This gives you a continuous score where positive values mean positive sentiment, negative mean negative:
library(tidytext) # 拆分文本为单词,匹配AFINN词典,按case_id求和得到每行情感分 sentiment_scores <- your_data %>% unnest_tokens(word, text) %>% # 把每行文本拆成单个单词 inner_join(get_sentiments("afinn")) %>% # 匹配情感词典 group_by(case_id) %>% summarise(sentiment_score = sum(value)) # 按案例ID求和得到该行总情感分 # 把情感分合并回原数据集 your_data_with_sentiment <- your_data %>% left_join(sentiment_scores, by = "case_id") %>% replace_na(list(sentiment_score = 0)) # 没有匹配到情感词的行填充0
选项B:获取情感分类(BING词典)
If you prefer categorical results (positive/negative/neutral) instead of numeric scores:
sentiment_categories <- your_data %>% unnest_tokens(word, text) %>% inner_join(get_sentiments("bing")) %>% # 匹配BING的正负词 group_by(case_id, sentiment) %>% count() %>% pivot_wider(names_from = sentiment, values_from = n, values_fill = 0) %>% # 根据正负词数量判断最终分类 mutate(sentiment_category = case_when( positive > negative ~ "positive", negative > positive ~ "negative", TRUE ~ "neutral" )) %>% select(case_id, sentiment_category) # 合并回原数据 your_data_with_sentiment <- your_data %>% left_join(sentiment_categories, by = "case_id") %>% replace_na(list(sentiment_category = "neutral"))
方法2:使用sentimentr(智能处理上下文)
If you want better handling of context (like negations: "not happy" gets recognized as negative, which basic word-based methods miss), sentimentr is perfect—it calculates sentiment directly on full lines of text without splitting words:
library(sentimentr) # 直接给每行文本计算情感分数,自动关联case_id your_data_with_sentiment <- your_data %>% mutate( sentiment_score = sentiment_by(text, by = case_id)$ave_sentiment, # 可选:把分数转成分类 sentiment_category = case_when( sentiment_score > 0 ~ "positive", sentiment_score < 0 ~ "negative", TRUE ~ "neutral" ) )
关键注意事项
- 数据导入: When loading your BI data, make sure
case_idis treated as a character (not numeric) to avoid merging issues. Useread_csv("your_data.csv", col_types = cols(case_id = col_character()))if needed. - 缺失值: Rows with no sentiment-related words will get
NAscores—usereplace_na()to fill these with 0 (for scores) or "neutral" (for categories) as shown above. - 词典选择: AFINN gives nuanced numeric scores, BING is simple for binary classification, and
sentimentrhandles context best. Pick the one that fits your use case!
内容的提问来源于stack exchange,提问作者Josh

