如何在R语言中移除字符串列中的指定常见后缀集合?
Remove Common Company Suffixes from String Endings
Got it, let's work through this solution to strip those top 100 company suffixes from the end of your db_name$name strings, while leaving the rest of the content intact.
Step 1: Fix the Regex Pattern (Critical!)
Your initial code builds a regex foundation, but it misses two key details to make it reliable:
- Escaping special characters: Some suffixes (like
Co.or& Co) contain regex-reserved characters (.,&, etc.) that need to be escaped so they match literally, not as wildcards. - Anchoring to the end: We need to ensure the regex only targets suffixes at the very end of the string, plus account for any leading space before the suffix (e.g., "Acme Corp" has a space before "Corp").
Here's the adjusted code to build a robust regex pattern:
library(stringr) # Extract the last word as you originally did db_name$ending <- word(db_name$name, -1) # Get the top 100 most frequent suffixes db_end_count <- data.frame(table(db_name$ending)) %>% arrange(desc(Freq)) %>% filter(row_number() <= 100) # Escape regex special characters in suffixes, then build the final pattern db_end_escaped <- str_escape(db_end_count$Var1) db_end_regex <- str_c("\\s?(", str_c(db_end_escaped, collapse = "|"), ")$")
Step 2: Remove the Suffixes Cleanly
Now use str_remove() to strip the suffix pattern from the end of each string. We'll add a safety check to avoid turning strings that are only a suffix into empty values:
# Remove suffixes, keep the original string if no valid suffix is found at the end db_name$clean_name <- ifelse( str_detect(db_name$name, db_end_regex), str_remove(db_name$name, db_end_regex), db_name$name )
Breakdown of the Solution
str_escape(): Handles tricky characters like.or+in suffixes, ensuring they're treated as literal text instead of regex commands.\\s?: Matches an optional space before the suffix (adjust or remove this if your suffixes are always attached without a space, like "AcmeInc").(...)$: Groups the suffix options and anchors the match to the end of the string, so we never accidentally remove suffixes that appear in the middle of a company name.- The
ifelse()check: Prevents edge cases where a string is exactly one of the suffixes (e.g., "Inc") from becoming an empty string.
Example Output
| Original Name | Clean Name |
|---|---|
| "Bright Tech Solutions Inc" | "Bright Tech Solutions" |
| "European Manufacturing GmbH" | "European Manufacturing" |
| "Just Co" | "Just" |
| "Ltd" | "Ltd" |
内容的提问来源于stack exchange,提问作者Nicholas John Pongratz
相关产品推荐
相关产品推荐

