Python re模块中如何使用\p{P}与\p{S}匹配标点和符号?
问题原因
Python 内置的 re 库原生不支持 Unicode 属性转义语法(即你用到的 \p{P}、\p{S} 这类写法),因此直接使用会抛出无效转义的正则错误。
解决方案
方案1:使用第三方 regex 库(兼容原写法)
该第三方库完全支持 \p{P}、\p{S} 语法,用法和内置 re 库基本一致:
- 安装命令:
pip install regex - 使用示例:
import regex pattern = r"[\p{P}\p{S}]" # 匹配测试 test_str = "测试字符串!@#,。" result = regex.findall(pattern, test_str) # 返回结果:['!', '@', '#', ',', '。']
方案2:使用标准库实现(无需安装第三方依赖)
有两种实现方式可以选:
- 正则表达式硬匹配
直接覆盖常用的中英文标点、特殊符号,写法如下:
import re pattern = r"[!\"#$%&'()*+,\-./:;<=>?@\[\]^_`{|}~!?,。、;:‘’“”()【】《》「」『』…—·¥……()]" test_str = "测试字符串!@#,。" result = re.findall(pattern, test_str)
- 结合
unicodedata标准库精准匹配
这种方式和\p{P}、\p{S}匹配逻辑完全一致,通过判断字符的 Unicode 类别即可:
import unicodedata def is_punct_or_symbol(char): # Unicode类别中P开头为标点,S开头为符号 return unicodedata.category(char).startswith(('P', 'S')) test_str = "测试字符串!@#,。" result = [c for c in test_str if is_punct_or_symbol(c)]
内容的提问来源于stack exchange,提问作者user1444989
相关产品推荐
相关产品推荐

