如何基于tm包或DocumentTermMatrix移除含3个及以上连续重复字母的词?
Hey there! Let's tackle this problem of stripping out those meaningless words with 3+ consecutive repeated letters from your text corpus. You can easily add this rule to your existing tm workflow using a custom function with tm_map—here's how to make it work:
Step 1: Define the Custom Removal Function
First, create a function that uses regex to spot and eliminate any word containing 3 or more identical consecutive letters. The regex targets whole words where at least one letter repeats 3+ times in a row:
remove_repeated_words <- function(text) { # Regex breakdown: # \\b = word boundary (ensures we target full words) # \\w* = any number of word characters before the repeated sequence # ([a-z])\\1{2,} = capture a lowercase letter, match it 2+ more times (total 3+) # \\w* = any number of word characters after the repeated sequence # \\b = word boundary gsub("\\b\\w*([a-z])\\1{2,}\\w*\\b", "", text) }
Step 2: Integrate into Your Preprocessing Pipeline
Insert this custom tm_map call into your workflow at the right spot—after converting to lowercase (so the regex matches consistently) and before removing stopwords. Here’s your updated full preprocessing code:
# Load required library library(tm) # Your sample input data text = c('apple', 'banana', 'orange', 'travelling', 'esteem', 'woooo','awwwwwwww','waaaaakakakakaka') # Create corpus tm_test <- VCorpus(VectorSource(text)) # Your existing preprocessing steps + new repeated word removal tm_test = tm_map(tm_test, removePunctuation) for (i in seq(tm_test)) { tm_test[[i]] = gsub("/", " ", tm_test[[i]]) tm_test[[i]] = gsub("@", " ", tm_test[[i]]) tm_test[[i]] = gsub("\\|", " ", tm_test[[i]]) } tm_test = tm_map(tm_test, removeNumbers) tm_test = tm_map(tm_test, tolower) # Add the new custom step here! tm_test = tm_map(tm_test, content_transformer(remove_repeated_words)) tm_test = tm_map(tm_test, PlainTextDocument) tm_test = tm_map(tm_test, removeWords, stopwords("english")) tm_test = tm_map(tm_test, PlainTextDocument) # tm_test = tm_map(tm_test, stemDocument) # tm_test = tm_map(tm_test, PlainTextDocument) tm_test = tm_map(tm_test, stripWhitespace) tm_test = tm_map(tm_test, PlainTextDocument) # Generate DocumentTermMatrix dtm_test = DocumentTermMatrix(tm_test)
Step 3: Check the Cleaned Result
If you extract the processed text from the corpus, you’ll get exactly the output you’re aiming for:
# Extract cleaned text and filter out empty strings cleaned_text <- sapply(tm_test, content) cleaned_text <- cleaned_text[cleaned_text != ""] print(cleaned_text) # Output: # [1] "apple" "banana" "orange" "travelling" "esteem"
A quick heads-up: We use content_transformer() with tm_map because our function operates directly on the text content—this ensures the tm package handles it correctly instead of treating it as a metadata operation.
内容的提问来源于stack exchange,提问作者recon

