You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何移除字符串中连续大写单词(含后续标点)并保留指定单词?

Solution for Removing Consecutive Uppercase Word Groups

Got it, let's fix this problem properly. The key here is to target only sequences of 2+ pure uppercase words (and any trailing punctuation after the last word in the sequence), while leaving single uppercase words and mixed-case/number words untouched.

Step 1: The Right Regular Expression

We need a regex pattern that specifically matches consecutive uppercase word clusters. Here's the pattern we'll use:

pattern <- "\\b[:upper:]+(\\s+[:upper:]+)+\\W*"

Let's break down what each part does:

  • \\b[:upper:]+: Matches a single pure uppercase word (the word boundary \\b ensures we don't match partial words)
  • (\\s+[:upper:]+)+: This captures one or more instances of "space + uppercase word" — this is what ensures we're targeting groups of 2 or more consecutive uppercase words
  • \\W*: Matches any trailing non-word characters (like ?, ., !) that come right after the final uppercase word in the cluster, so we can remove those too

Step 2: Implement the Fix

Let's apply this to your example string using stringr::str_remove_all():

library(stringr)

string <- "Lorem ipsum DOLOR SIT AMET? consectetuer adipiscing elit. Morbi gravida libero NEC velit. Morbi scelerisque luctus velit. ETIAM-123 dui sem, fermentum vitae, SAGITTIS ID? malesuada in, quam. Proin mattis lacinia justo. Vestibulum facilisis auctor urna. Aliquam IN LOREM SIT amet leo accumsan"

clean_string <- str_remove_all(string, pattern)

# Check the result
cat(clean_string)

Step 3: Verify the Result

Running the code above will give you this cleaned string:

Lorem ipsum consectetuer adipiscing elit. Morbi gravida libero NEC velit. Morbi scelerisque luctus velit. ETIAM-123 dui sem, fermentum vitae, malesuada in, quam. Proin mattis lacinia justo. Vestibulum facilisis auctor urna. Aliquam amet leo accumsan

Notice that:

  • The clusters DOLOR SIT AMET?, SAGITTIS ID?, and IN LOREM SIT are all removed (including the trailing ? in the first two)
  • Single uppercase word NEC is preserved
  • Mixed word ETIAM-123 is completely preserved (since it's not a pure uppercase word, our regex doesn't touch it)

Why Your Original Code Failed

Your original pattern \\b[:upper:]+\\b matches all pure uppercase words, regardless of whether they're alone or in a group. That's why it removed NEC and chopped ETIAM-123 down to -123 — it only matched the ETIAM part. Our targeted regex avoids this by only going after clusters of 2+ pure uppercase words.

内容的提问来源于stack exchange,提问作者gebo-aa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 19:27:57