You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言文本分析中如何去除单词末尾s及缩写字母?

Fixing Trailing 's' and Abbreviation Removal in R Text Analysis

Hey there! Let's work through your problem of removing trailing "s" (like plural endings) and handling abbreviation-related text in your R pipeline. Your current preprocessing covers the basics well, but we need to add targeted steps to address those suffixes specifically.

Step 1: Handle Possessive 's' First

Before tackling plural "s", let's strip out possessive endings like "John's" or "company's"—these are easy to target with a regex that matches word-boundary trailing 's:

# Remove possessive 's' (e.g., "Trump's" → "Trump")
x <- gsub("'s\\b", "", x)

Step 2: Remove Plural Trailing 's'

To strip plural "s" from nouns (this will also affect verb third-person singular forms, which is often acceptable in text analysis unless you need to preserve tense), use a regex that captures the preceding character and replaces the trailing "s":

# Remove trailing 's' from words (e.g., "cats" → "cat", "runs" → "run")
x <- gsub("([a-zA-Z])s\\b", "\\1", x)

If you need to avoid modifying verb forms, you'd need to add part-of-speech tagging (e.g., with the spacyr package) to only target nouns, but that adds complexity. For most text analysis use cases, the above regex works great.

Step 3: Target Specific Abbreviations

For other common abbreviations (like "dept" → "department", "govt" → "government"), you can use a custom replacement dictionary. Here's how to do it with stringr:

# Define your abbreviation key
abbrev_key <- c(
  "\\bdept\\b" = "department",
  "\\bgovt\\b" = "government",
  "\\betc\\b" = "etcetera",
  "\\binfo\\b" = "information"
)

# Replace abbreviations
x <- stringr::str_replace_all(x, abbrev_key)

Adjust the key to match the abbreviations you see in your dataset.

Updated Full Preprocessing Pipeline

Here's how these steps fit into your existing code:

x <- demtweets$Tweet
x <- paste(unlist(x), collapse =" ")
x <- stringi::stri_trans_general(x, "latin-ascii")
x <- gsub(" '[A-z] ", " ", x)
x <- gsub("&amp;amp", "", x)
x <- gsub("(RT|via)((?:\\b\\W*@\\w+)+)", "", x)
x <- gsub("@\\w+", "", x)
x <- gsub("[[:punct:]]", "", x)
x <- gsub("[[:digit:]]", "", x)
x <- gsub("http\\w+", "", x)
x <- gsub("[ \t]{2,}", "", x)
x <- gsub("^\\s+|\\s+$", "", x)

# New steps for possessives, plural s, and abbreviations
x <- gsub("'s\\b", "", x)
x <- gsub("([a-zA-Z])s\\b", "\\1", x)
abbrev_key <- c(
  "\\bdept\\b" = "department",
  "\\bgovt\\b" = "government",
  "\\betc\\b" = "etcetera"
)
x <- stringr::str_replace_all(x, abbrev_key)

# Continue with contraction replacement and dfm creation
x <- replace_contraction(x, contraction.key = lexicon::key_contractions, ignore.case = TRUE)
x <- replace_contraction(x, contraction = qdapDictionaries::contractions, replace = NULL, ignore.case = TRUE)

xdfm <- dfm(x, stem = F, remove_punct = T, tolower = T, remove_twitter = T, remove_numbers = TRUE, remove = c(stopwords("english"), "http","https","rt", "t.co"))
textplot_wordcloud(xdfm, min_count = 6, random_order = FALSE, rotation = .25, color = RColorBrewer::brewer.pal(8, "Dark2"))
topfeatures(xdfm, 100)

Notes

  • If you notice over-replacement (e.g., words like "bus" turning into "bu"), tweak the regex to exclude words where the last character before "s" is a vowel: gsub("([^aeiouAEIOU])s\\b", "\\1", x)—this will leave words like "bus" intact while still handling "cats" or "dogs".
  • For more advanced abbreviation handling, check out the textclean package, which has a built-in replace_abbreviation() function covering hundreds of common abbreviations.

内容的提问来源于stack exchange,提问作者Abdul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:55:55