使用re.sub()移除特殊字符时,replace()替换功能失效问题排查
问题描述
尝试使用re.sub()函数移除特殊字符,但调用该函数后,此前replace()的字符串替换效果全部消失。
实现代码
import re import pandas as pd from IPython.display import display tabela = pd.read_excel("tst.xlsx") (tabela[['nome', 'mensagem', 'arquivo', 'telefone']]) for linha in tabela.index: nome = tabela.loc[linha, "nome"] mensagem = tabela.loc[linha, "mensagem"] acordo = tabela.loc[linha, "acordo"] telefone = tabela.loc[linha, "telefone"] texto = mensagem.replace("fulano", nome) texto = texto.replace( "value", acordo) texto = texto.replace( "phone", telefone) texto = re.sub(r"[!!@#$%¨&*()_?',;.]", '', telefone) print(texto)
实际输出
11
预期输出
thyago R$200 11
解决方案
问题出在最后一行的赋值逻辑:你把texto重新赋值为re.sub()处理telefone的结果,直接覆盖了之前三次replace()生成的内容,导致之前的替换效果全部丢失。
修正方案1:保留完整替换结果后清理特殊字符
将re.sub()的处理对象改为前面已经完成替换的texto:
import re import pandas as pd from IPython.display import display tabela = pd.read_excel("tst.xlsx") (tabela[['nome', 'mensagem', 'arquivo', 'telefone']]) for linha in tabela.index: nome = tabela.loc[linha, "nome"] mensagem = tabela.loc[linha, "mensagem"] acordo = tabela.loc[linha, "acordo"] telefone = tabela.loc[linha, "telefone"] texto = mensagem.replace("fulano", nome) texto = texto.replace( "value", acordo) texto = texto.replace( "phone", telefone) # 处理对象改为已完成替换的texto,保留之前的替换结果 texto = re.sub(r"[!!@#$%¨&*()_?',;.]", '', texto) print(texto)
修正方案2:提前清理手机号再替换
如果仅需要移除telefone中的特殊字符再替换到文本中,可以提前处理手机号:
import re import pandas as pd from IPython.display import display tabela = pd.read_excel("tst.xlsx") (tabela[['nome', 'mensagem', 'arquivo', 'telefone']]) for linha in tabela.index: nome = tabela.loc[linha, "nome"] mensagem = tabela.loc[linha, "mensagem"] acordo = tabela.loc[linha, "acordo"] telefone = tabela.loc[linha, "telefone"] # 先清理手机号的特殊字符 telefone_limpo = re.sub(r"[!!@#$%¨&*()_?',;.]", '', telefone) texto = mensagem.replace("fulano", nome) texto = texto.replace( "value", acordo) texto = texto.replace( "phone", telefone_limpo) print(texto)
内容的提问来源于stack exchange,提问作者Thyago
相关产品推荐
相关产品推荐

