Python正则表达式分组问题:Pandas Series替换失败求助
问题分析与解决方法
你的正则代码存在三个核心错误:
未启用正则匹配模式
Pandas的Series.replace()默认参数regex=False,会把第一个参数当作普通字符串做字面匹配,而非正则表达式。你的原字符串中数字是变化的,字面匹配根本找不到完全一致的内容,必须添加regex=True才会按正则规则处理。正则语法错误
你用{}包裹数字匹配规则,这是误用。正则里{}是量词(比如\d{2}匹配两位数字),不是用来包裹表达式的。正确的数字匹配写法是直接写\d+(\.\d+)?(匹配整数或小数),不需要加{}。正则逻辑冗余
你没必要精确匹配后面每一组expected: x || is: y,因为目标是保留前面固定的前缀,直接匹配从##到末尾的所有内容即可,写法更简洁高效。
正确解法(两种可选)
解法1:移除冗余后缀(推荐)
直接匹配从第一个##到最后一个##的所有内容,替换为空:
import pandas as pd m = pd.Series(['expected != is --> found missing lices ## expected: 2.25 || is: 4.5 || expected: 3 || is: 2 ##','expected != is --> found missing lices ## expected: 3.35 || is: 5.5 || expected: 3 || is: 3 ##', 'expected != is --> found missing lices ## expected: 2.25 || is: 4.5 || expected: 3 || is: 2 ##']) # 替换从##开始到结尾的所有内容 m = m.replace(r'##.*##', '', regex=True)
解法2:修正你的正则写法
保留你原有的匹配逻辑,修正语法并启用正则模式:
m = m.replace( r'expected != is --> found missing lices ## expected: \d+(\.\d+)? || is: \d+(\.\d+)? || expected: \d+ || is: \d+ ##', 'expected != is --> found missing lices', regex=True )
解法3:直接替换为目标字符串(如果所有元素都符合前缀规则)
如果每个元素都以目标字符串开头,也可以直接替换整个内容:
m = m.str.replace(r'.*', 'expected != is --> found missing lices', regex=True)
内容的提问来源于stack exchange,提问作者shir13
相关产品推荐
相关产品推荐

