如何用gsub移除字符串中括号内的英文字符(适配多括号场景)
Got it, let's sort out this multi-parentheses issue for you!
The Problem with Your Original Code
Your initial regex gsub('\(.*?[a-zA-Z]\)', '', x) works fine for single parentheses, but it fails with multiple pairs because it doesn't properly target complete parentheses pairs that contain English characters. Instead, the non-greedy .*? will skip over parentheses without English text and keep matching until it hits the first English character in a later pair—this ends up deleting way more content than you want.
The Solution: Target Complete Parentheses Pairs with English Text
We need a regex that specifically matches entire parentheses pairs where there's at least one English character inside. Here's the adjusted code:
# Your multi-parentheses input x <- "Определение полиморфизма A1648G (замена аденина на гуанин в позиции 1648) в гене гликопротеина la (GpIA)" # Remove only parentheses pairs containing English characters result <- gsub('\\([^)]*[a-zA-Z][^)]*\\)', '', x) # Print the result cat(result)
Output:
Определение полиморфизма A1648G (замена аденина на гуанин в позиции 1648) в гене гликопротеина la
Clean Up Extra Spaces
If you want to get rid of the trailing space (or any extra spaces around removed parentheses), add space matching to the regex and use trimws():
result_clean <- trimws(gsub('\\s*\\([^)]*[a-zA-Z][^)]*\\)\\s*', ' ', x)) cat(result_clean)
Clean Output:
Определение полиморфизма A1648G (замена аденина на гуанин в позиции 1648) в гене гликопротеина la
Regex Breakdown
Let's break down why this works:
\\(: Matches the opening parenthesis (escaped because parentheses are special regex characters)[^)]*: Matches any number of characters except a closing parenthesis—this ensures we don't cross into other parentheses pairs[a-zA-Z]: Ensures we only target parentheses that contain at least one English character[^)]*: Matches any remaining characters inside the parentheses pair\\): Matches the closing parenthesis\\s*: Optional leading/trailing spaces to clean up whitespace after removal
Verify with Single Parentheses
This code also works for your original single-parentheses case:
x_single <- "Бактериологическое исследование отделяемого слизистой оболочки прямой кишки на гонококк (Nesseria gonorrheeae)" result_single <- trimws(gsub('\\s*\\([^)]*[a-zA-Z][^)]*\\)\\s*', ' ', x_single)) cat(result_single)
Output:
Бактериологическое исследование отделяемого слизистой оболочки прямой кишки на гонококк
内容的提问来源于stack exchange,提问作者john

