You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用gsub移除字符串中括号内的英文字符(适配多括号场景)

Fixing Multi-Parentheses Removal of English Characters in Strings

Got it, let's sort out this multi-parentheses issue for you!

The Problem with Your Original Code

Your initial regex gsub('\(.*?[a-zA-Z]\)', '', x) works fine for single parentheses, but it fails with multiple pairs because it doesn't properly target complete parentheses pairs that contain English characters. Instead, the non-greedy .*? will skip over parentheses without English text and keep matching until it hits the first English character in a later pair—this ends up deleting way more content than you want.

The Solution: Target Complete Parentheses Pairs with English Text

We need a regex that specifically matches entire parentheses pairs where there's at least one English character inside. Here's the adjusted code:

# Your multi-parentheses input
x <- "Определение полиморфизма A1648G (замена аденина на гуанин в позиции 1648) в гене гликопротеина la (GpIA)"

# Remove only parentheses pairs containing English characters
result <- gsub('\\([^)]*[a-zA-Z][^)]*\\)', '', x)

# Print the result
cat(result)

Output:

Определение полиморфизма A1648G (замена аденина на гуанин в позиции 1648) в гене гликопротеина la 

Clean Up Extra Spaces

If you want to get rid of the trailing space (or any extra spaces around removed parentheses), add space matching to the regex and use trimws():

result_clean <- trimws(gsub('\\s*\\([^)]*[a-zA-Z][^)]*\\)\\s*', ' ', x))
cat(result_clean)

Clean Output:

Определение полиморфизма A1648G (замена аденина на гуанин в позиции 1648) в гене гликопротеина la

Regex Breakdown

Let's break down why this works:

  • \\(: Matches the opening parenthesis (escaped because parentheses are special regex characters)
  • [^)]*: Matches any number of characters except a closing parenthesis—this ensures we don't cross into other parentheses pairs
  • [a-zA-Z]: Ensures we only target parentheses that contain at least one English character
  • [^)]*: Matches any remaining characters inside the parentheses pair
  • \\): Matches the closing parenthesis
  • \\s*: Optional leading/trailing spaces to clean up whitespace after removal

Verify with Single Parentheses

This code also works for your original single-parentheses case:

x_single <- "Бактериологическое исследование отделяемого слизистой оболочки прямой кишки на гонококк (Nesseria gonorrheeae)"
result_single <- trimws(gsub('\\s*\\([^)]*[a-zA-Z][^)]*\\)\\s*', ' ', x_single))
cat(result_single)

Output:

Бактериологическое исследование отделяемого слизистой оболочки прямой кишки на гонококк

内容的提问来源于stack exchange,提问作者john

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:32:55