R语言DataFrame文本列按停用词列表替换指定词汇且保留词序的实现问题
R 文本停用词替换实现方案
你的代码出错的核心原因是str_replace的第一个输入参数被你写成了逻辑判断表达式,返回的是布尔值而非待处理的单词文本,所以最终得到的是逻辑向量而非替换后的单词列表,同时冗余的循环和多层lapply也增加了逻辑混乱。
以下是满足需求的完整实现:
1. 基础数据定义
# 示例数据集 df <- structure(list(ID = 1:3, Text = c("there was not clostridium", "clostridium difficile positive", "test was OK but there was clostridium")), class = "data.frame", row.names = c(NA, -3L)) # 停用词正则模式(避免和基础函数stop重名,改名stop_pattern) stop_pattern <- paste0(c("was", "but", "there"), collapse = "|")
2. 处理代码
无需循环、无需merge类函数,单层遍历即可实现,自动保留原有词序:
library(tokenizers) library(stringr) # 生成分词列表 df$Words <- tokenize_words(df$Text, lowercase = TRUE) # 批量替换停用词 df$clean <- lapply(df$Words, function(word_vec) { str_replace(word_vec, pattern = stop_pattern, replacement = "REPLACED") }) # 可选:将列表列转为逗号分隔的字符串,方便打印查看 df$Words <- sapply(df$Words, paste, collapse = ", ") df$clean <- sapply(df$clean, paste, collapse = ", ")
3. 输出结果
打印df即可得到你预期的效果:
ID Text Words clean 1 1 there was not clostridium there, was, not, clostridium REPLACED, REPLACED, not, clostridium 2 2 clostridium difficile positive clostridium, difficile, positive clostridium, difficile, positive 3 3 test was OK but there was clostridium test, was, ok, but, there, was, clostridium test, REPLACED, ok, REPLACED, REPLACED, REPLACED, clostridium
内容的提问来源于stack exchange,提问作者onhalu
相关产品推荐
相关产品推荐

