You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于R语言利用现有词典匹配分析调研评论的技术问询

解决R语言中调研评论的多词典匹配问题

Hey there! I’ve got you covered for this dictionary-based survey comment analysis. Let’s break this down into simple, actionable steps with code examples you can adapt right away.

1. 先准备示例数据

First, let's set up some sample survey comments and dictionaries to work with. I’ll use relatable categories like positive sentiment, negative sentiment, and service-related terms:

# 示例调研评论
survey_comments <- c(
  "The staff was incredibly friendly and helpful!",
  "Waited 2 hours with no update, terrible experience.",
  "Just okay.",
  "Loved the quick checkout process!",
  "Food was cold, won't be coming back."
)

# 示例词典
positive_dict <- c("friendly", "helpful", "loved", "quick")
negative_dict <- c("terrible", "cold", "won't", "wait")
service_dict <- c("staff", "checkout", "update", "experience")

2. 编写匹配函数

We’ll create a helper function that checks if a comment contains any term from a given dictionary. This function will handle case insensitivity (so "Friendly" matches "friendly") and full-word matches (to avoid false positives like "wait" matching "waited" if you don’t want that—we’ll adjust for that too).

# 定义匹配函数:检查评论是否包含词典中的任意词汇
check_dictionary_match <- function(comment, dict, full_word = TRUE) {
  # 构建正则表达式:如果是完整单词匹配,加上词边界\\b
  pattern <- if(full_word) {
    paste0("\\b", paste(dict, collapse = "\\b|\\b"), "\\b")
  } else {
    paste(dict, collapse = "|")
  }
  # 检查匹配,忽略大小写
  grepl(pattern, comment, ignore.case = TRUE)
}

3. 生成结果DataFrame

Now we’ll apply this function to each dictionary and combine everything into a single dataframe where the first column is the original comment, followed by True/False columns for each dictionary:

# 创建结果数据框
result_df <- data.frame(
  Survey_Comment = survey_comments,
  Positive_Match = sapply(survey_comments, check_dictionary_match, dict = positive_dict),
  Negative_Match = sapply(survey_comments, check_dictionary_match, dict = negative_dict),
  Service_Match = sapply(survey_comments, check_dictionary_match, dict = service_dict),
  stringsAsFactors = FALSE
)

# 查看结果
print(result_df)

4. 输出示例

When you run the code above, you’ll get a dataframe that looks like this:

Survey_Comment Positive_Match Negative_Match Service_Match
1 The staff was incredibly friendly and helpful!           TRUE          FALSE          TRUE
2   Waited 2 hours with no update, terrible experience.          FALSE           TRUE          TRUE
3                                      Just okay.          FALSE          FALSE         FALSE
4                Loved the quick checkout process!           TRUE          FALSE          TRUE
5                Food was cold, won't be coming back.          FALSE           TRUE         FALSE

额外调整建议

  • 如果不需要完整单词匹配(比如想让 "waited" 匹配 "wait"),只需要在调用函数时设置 full_word = FALSE。
  • 如果你有很多词典,可以把它们放进一个列表里,然后用 purrr::map_dfc 来批量生成列,这样更高效:
    library(purrr)
    
    # 把词典放进列表
    dict_list <- list(
      Positive = positive_dict,
      Negative = negative_dict,
      Service = service_dict
    )
    
    # 批量生成匹配列
    match_columns <- map_dfc(dict_list, ~sapply(survey_comments, check_dictionary_match, dict = .x))
    
    # 合并到结果
    result_df <- bind_cols(data.frame(Survey_Comment = survey_comments), match_columns)
    

内容的提问来源于stack exchange,提问作者thewrightowns

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:20:53