在R语言文本分析中如何去除单词末尾s及缩写字母?
Hey there! Let's work through your problem of removing trailing "s" (like plural endings) and handling abbreviation-related text in your R pipeline. Your current preprocessing covers the basics well, but we need to add targeted steps to address those suffixes specifically.
Step 1: Handle Possessive 's' First
Before tackling plural "s", let's strip out possessive endings like "John's" or "company's"—these are easy to target with a regex that matches word-boundary trailing 's:
# Remove possessive 's' (e.g., "Trump's" → "Trump") x <- gsub("'s\\b", "", x)
Step 2: Remove Plural Trailing 's'
To strip plural "s" from nouns (this will also affect verb third-person singular forms, which is often acceptable in text analysis unless you need to preserve tense), use a regex that captures the preceding character and replaces the trailing "s":
# Remove trailing 's' from words (e.g., "cats" → "cat", "runs" → "run") x <- gsub("([a-zA-Z])s\\b", "\\1", x)
If you need to avoid modifying verb forms, you'd need to add part-of-speech tagging (e.g., with the spacyr package) to only target nouns, but that adds complexity. For most text analysis use cases, the above regex works great.
Step 3: Target Specific Abbreviations
For other common abbreviations (like "dept" → "department", "govt" → "government"), you can use a custom replacement dictionary. Here's how to do it with stringr:
# Define your abbreviation key abbrev_key <- c( "\\bdept\\b" = "department", "\\bgovt\\b" = "government", "\\betc\\b" = "etcetera", "\\binfo\\b" = "information" ) # Replace abbreviations x <- stringr::str_replace_all(x, abbrev_key)
Adjust the key to match the abbreviations you see in your dataset.
Updated Full Preprocessing Pipeline
Here's how these steps fit into your existing code:
x <- demtweets$Tweet x <- paste(unlist(x), collapse =" ") x <- stringi::stri_trans_general(x, "latin-ascii") x <- gsub(" '[A-z] ", " ", x) x <- gsub("&amp", "", x) x <- gsub("(RT|via)((?:\\b\\W*@\\w+)+)", "", x) x <- gsub("@\\w+", "", x) x <- gsub("[[:punct:]]", "", x) x <- gsub("[[:digit:]]", "", x) x <- gsub("http\\w+", "", x) x <- gsub("[ \t]{2,}", "", x) x <- gsub("^\\s+|\\s+$", "", x) # New steps for possessives, plural s, and abbreviations x <- gsub("'s\\b", "", x) x <- gsub("([a-zA-Z])s\\b", "\\1", x) abbrev_key <- c( "\\bdept\\b" = "department", "\\bgovt\\b" = "government", "\\betc\\b" = "etcetera" ) x <- stringr::str_replace_all(x, abbrev_key) # Continue with contraction replacement and dfm creation x <- replace_contraction(x, contraction.key = lexicon::key_contractions, ignore.case = TRUE) x <- replace_contraction(x, contraction = qdapDictionaries::contractions, replace = NULL, ignore.case = TRUE) xdfm <- dfm(x, stem = F, remove_punct = T, tolower = T, remove_twitter = T, remove_numbers = TRUE, remove = c(stopwords("english"), "http","https","rt", "t.co")) textplot_wordcloud(xdfm, min_count = 6, random_order = FALSE, rotation = .25, color = RColorBrewer::brewer.pal(8, "Dark2")) topfeatures(xdfm, 100)
Notes
- If you notice over-replacement (e.g., words like "bus" turning into "bu"), tweak the regex to exclude words where the last character before "s" is a vowel:
gsub("([^aeiouAEIOU])s\\b", "\\1", x)—this will leave words like "bus" intact while still handling "cats" or "dogs". - For more advanced abbreviation handling, check out the
textcleanpackage, which has a built-inreplace_abbreviation()function covering hundreds of common abbreviations.
内容的提问来源于stack exchange,提问作者Abdul

