You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则移除未被字母数字包围的非字母数字字符?

解决方案

要实现移除字符串中所有非字母数字字符,但保留被字母数字字符包围的非字母数字字符,可以使用Python的re模块,通过匹配并删除不符合保留条件的字符来完成。

正则表达式与代码实现

import re

def clean_target_string(input_str):
    # 匹配需要删除的非字母数字/非空格字符:要么左侧无字母数字,要么右侧无字母数字
    pattern = r'(?<![a-zA-Z0-9])[^a-zA-Z0-9\s]|[^a-zA-Z0-9\s](?![a-zA-Z0-9])'
    return re.sub(pattern, '', input_str)

# 测试示例
print(clean_target_string("it's"))  # 输出: it's
print(clean_target_string("its."))  # 输出: its
test_input = "It's a: beautiful day? I'm =sure it is. The coca-cola (is frozen right?"
print(clean_target_string(test_input))  
# 输出: It's a beautiful day I'm sure it is The coca-cola is frozen right

正则逻辑说明

正则表达式分为两个分支,分别匹配需要删除的字符:

  • (?<![a-zA-Z0-9])[^a-zA-Z0-9\s]:匹配左侧没有字母/数字的非字母数字、非空格字符(比如示例中的=、(、:)
  • [^a-zA-Z0-9\s](?![a-zA-Z0-9]):匹配右侧没有字母/数字的非字母数字、非空格字符(比如示例中的.、?)

而像it's里的'、coca-cola里的-这类被字母数字左右包围的字符,不会被上述分支匹配,因此会被保留。

内容的提问来源于stack exchange,提问作者BoomBoxBoy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 17:15:53