如何使用re.sub移除字符串中的特定符号(含Unicode字符)
问题:如何移除文本中的特殊象形字符?
原始文本及尝试的代码如下:
初始文本:
txt = 'model i love u take with u all the time in ur± ' #此处含中文象形字符
尝试直接匹配符号移除:
import re txt = re.sub("[±]", '', txt) print(txt) # 输出:'model i love u take with u all the time in ur\x9f\x93± \x9f\x9f\x98\x8e\x9f\x91\x84\x9f\x91\x9f\x92\x9f\x92\x9f\x92'
改用Unicode编码匹配移除:
txt = re.sub("[\x93±\x98\x8e\x91\x84\x91\x9f\x92]", '', txt) print(txt) # 输出:'model i love u take with u all the time in ur\x9f\x93± \x9f\x9f\x98\x8e\x9f\x91\x84\x9f\x91\x9f\x92\x9f\x92\x9f\x92'
两种方式均未生效,该如何移除这类符号?
解决方案
方法1:保留ASCII可打印字符
直接过滤所有非ASCII的可打印字符,只保留字母、数字、空格和常见符号:
import re txt = 'model i love u take with u all the time in ur± ' # 含特殊字符的原始文本 cleaned_txt = re.sub(r'[^\x20-\x7E]', '', txt) print(cleaned_txt) # 输出:'model i love u take with u all the time in ur '
解释:\x20-\x7E 覆盖了所有ASCII可打印字符(空格到波浪线),不在此范围内的字符都会被移除。
方法2:精准匹配Emoji/象形字符的Unicode范围
如果要针对性移除表情符号类象形字符,可匹配对应的Unicode区块:
import re txt = 'model i love u take with u all the time in ur± ' # 含特殊字符的原始文本 # 匹配常见表情、象形字符的Unicode范围 cleaned_txt = re.sub(r'[\U0001F600-\U0001F64F\U0001F300-\U0001F5FF\U0001F680-\U0001F6FF\U0001F700-\U0001F77F\U0001F780-\U0001F7FF\U0001F800-\U0001F8FF\U0001F900-\U0001F9FF\U0001FA00-\U0001FA6F\U0001FA70-\U0001FAFF\U00002702-\U000027B0\U000024C2-\U0001F251]+', '', txt) # 单独移除±符号 cleaned_txt = re.sub(r'±', '', cleaned_txt) print(cleaned_txt) # 输出:'model i love u take with u all the time in ur '
方法3:利用字符串isprintable()方法
遍历字符串,只保留可打印的ASCII字符:
txt = 'model i love u take with u all the time in ur± ' # 含特殊字符的原始文本 cleaned_txt = ''.join([c for c in txt if c.isprintable() and ord(c) <= 127]) print(cleaned_txt) # 输出:'model i love u take with u all the time in ur '
解释:isprintable() 判断字符是否可打印,ord(c) <=127 确保是ASCII字符,可过滤所有非ASCII特殊象形字符和不可打印字符。
内容的提问来源于stack exchange,提问作者senek
相关产品推荐
相关产品推荐

