You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于tm包或DocumentTermMatrix移除含3个及以上连续重复字母的词?

Removing Words with 3+ Consecutive Repeated Letters in Text Mining Workflow

Hey there! Let's tackle this problem of stripping out those meaningless words with 3+ consecutive repeated letters from your text corpus. You can easily add this rule to your existing tm workflow using a custom function with tm_map—here's how to make it work:

Step 1: Define the Custom Removal Function

First, create a function that uses regex to spot and eliminate any word containing 3 or more identical consecutive letters. The regex targets whole words where at least one letter repeats 3+ times in a row:

remove_repeated_words <- function(text) {
  # Regex breakdown:
  # \\b = word boundary (ensures we target full words)
  # \\w* = any number of word characters before the repeated sequence
  # ([a-z])\\1{2,} = capture a lowercase letter, match it 2+ more times (total 3+)
  # \\w* = any number of word characters after the repeated sequence
  # \\b = word boundary
  gsub("\\b\\w*([a-z])\\1{2,}\\w*\\b", "", text)
}

Step 2: Integrate into Your Preprocessing Pipeline

Insert this custom tm_map call into your workflow at the right spot—after converting to lowercase (so the regex matches consistently) and before removing stopwords. Here’s your updated full preprocessing code:

# Load required library
library(tm)

# Your sample input data
text = c('apple', 'banana', 'orange', 'travelling', 'esteem', 'woooo','awwwwwwww','waaaaakakakakaka')

# Create corpus
tm_test <- VCorpus(VectorSource(text))

# Your existing preprocessing steps + new repeated word removal
tm_test = tm_map(tm_test, removePunctuation)
for (i in seq(tm_test)) {
  tm_test[[i]] = gsub("/", " ", tm_test[[i]])
  tm_test[[i]] = gsub("@", " ", tm_test[[i]])
  tm_test[[i]] = gsub("\\|", " ", tm_test[[i]])
}
tm_test = tm_map(tm_test, removeNumbers)
tm_test = tm_map(tm_test, tolower)
# Add the new custom step here!
tm_test = tm_map(tm_test, content_transformer(remove_repeated_words))
tm_test = tm_map(tm_test, PlainTextDocument)
tm_test = tm_map(tm_test, removeWords, stopwords("english"))
tm_test = tm_map(tm_test, PlainTextDocument)
# tm_test = tm_map(tm_test, stemDocument)
# tm_test = tm_map(tm_test, PlainTextDocument)
tm_test = tm_map(tm_test, stripWhitespace)
tm_test = tm_map(tm_test, PlainTextDocument)

# Generate DocumentTermMatrix
dtm_test = DocumentTermMatrix(tm_test)

Step 3: Check the Cleaned Result

If you extract the processed text from the corpus, you’ll get exactly the output you’re aiming for:

# Extract cleaned text and filter out empty strings
cleaned_text <- sapply(tm_test, content)
cleaned_text <- cleaned_text[cleaned_text != ""]

print(cleaned_text)
# Output:
# [1] "apple"      "banana"     "orange"     "travelling" "esteem"

A quick heads-up: We use content_transformer() with tm_map because our function operates directly on the text content—this ensures the tm package handles it correctly instead of treating it as a metadata operation.

内容的提问来源于stack exchange,提问作者recon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:42:00