You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用正则表达式移除指定字符串及其对应标点符号?

解决方案

问题分析

你的正则表达式未限制匹配完整单词,且对目标词前后标点与空格的组合处理不够灵活,替换后还会产生多余连续空格,导致部分场景处理失效。

修正后的代码

import re

def remove_target_words(text):
    # 匹配完整的bro/girl单词,以及前后的标点、空格组合
    pattern = r'\s*[,.!?]*\b(bro|girl)\b[,.!?]*\s*'
    # 先将目标内容替换为单个空格
    processed_text = re.sub(pattern, ' ', text)
    # 合并多个连续空格为一个,再去除首尾空格
    processed_text = re.sub(r'\s+', ' ', processed_text).strip()
    return processed_text

# 测试示例
test_samples = [
    "What's up, bro? That's cool",
    "You go girl! Go get them",
    "bro! Let's hang out",
    "Hey girl, what's up?",
    "Don't call me bro"
]

for sample in test_samples:
    print(f"原句: {sample}")
    print(f"处理后: {remove_target_words(sample)}\n")

正则规则说明

  • \b: 单词边界,确保只匹配完整的bro或girl单词,避免误匹配brother、girlfriend这类词中的片段
  • [,.!?]*: 匹配目标词前后可选的标点符号,可根据需求添加更多标点类型
  • \s*: 匹配目标词前后可选的空白字符
  • 二次替换\s+是为了清理替换后产生的多余连续空格,保证输出格式整洁

内容的提问来源于stack exchange,提问作者bigboyben123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 05:52:43