如何移除字符串中连续大写单词(含后续标点)并保留指定单词?
Got it, let's fix this problem properly. The key here is to target only sequences of 2+ pure uppercase words (and any trailing punctuation after the last word in the sequence), while leaving single uppercase words and mixed-case/number words untouched.
Step 1: The Right Regular Expression
We need a regex pattern that specifically matches consecutive uppercase word clusters. Here's the pattern we'll use:
pattern <- "\\b[:upper:]+(\\s+[:upper:]+)+\\W*"
Let's break down what each part does:
\\b[:upper:]+: Matches a single pure uppercase word (the word boundary\\bensures we don't match partial words)(\\s+[:upper:]+)+: This captures one or more instances of "space + uppercase word" — this is what ensures we're targeting groups of 2 or more consecutive uppercase words\\W*: Matches any trailing non-word characters (like?,.,!) that come right after the final uppercase word in the cluster, so we can remove those too
Step 2: Implement the Fix
Let's apply this to your example string using stringr::str_remove_all():
library(stringr) string <- "Lorem ipsum DOLOR SIT AMET? consectetuer adipiscing elit. Morbi gravida libero NEC velit. Morbi scelerisque luctus velit. ETIAM-123 dui sem, fermentum vitae, SAGITTIS ID? malesuada in, quam. Proin mattis lacinia justo. Vestibulum facilisis auctor urna. Aliquam IN LOREM SIT amet leo accumsan" clean_string <- str_remove_all(string, pattern) # Check the result cat(clean_string)
Step 3: Verify the Result
Running the code above will give you this cleaned string:
Lorem ipsum consectetuer adipiscing elit. Morbi gravida libero NEC velit. Morbi scelerisque luctus velit. ETIAM-123 dui sem, fermentum vitae, malesuada in, quam. Proin mattis lacinia justo. Vestibulum facilisis auctor urna. Aliquam amet leo accumsan
Notice that:
- The clusters
DOLOR SIT AMET?,SAGITTIS ID?, andIN LOREM SITare all removed (including the trailing?in the first two) - Single uppercase word
NECis preserved - Mixed word
ETIAM-123is completely preserved (since it's not a pure uppercase word, our regex doesn't touch it)
Why Your Original Code Failed
Your original pattern \\b[:upper:]+\\b matches all pure uppercase words, regardless of whether they're alone or in a group. That's why it removed NEC and chopped ETIAM-123 down to -123 — it only matched the ETIAM part. Our targeted regex avoids this by only going after clusters of 2+ pure uppercase words.
内容的提问来源于stack exchange,提问作者gebo-aa

