为什么正则分组处理字符串列表时,有时视为字符串有时返回匹配对象?
问题成因
- 提取结果未赋值:代码中仅执行了
number.group(0)、asker.group(1)这类操作,但没有将返回的字符串赋值给对应变量,导致变量还是原始的re.Match对象,输出时就会显示Match对象的格式内容。 - 匹配范围错误:直接将嵌套列表
text_pos整体转为字符串匹配,配合re.DOTALL贪婪匹配规则,会跨列表元素捕获到字符串里的多余符号(比如列表转义的引号、逗号、换行符等),无法精准提取目标内容。 - 正则逻辑与遍历逻辑不匹配:后续修改的逐字符串遍历逻辑,和需要跨元素匹配的
aaa正则规则不兼容,比如K开头的编号需要拼接当前元素和下一个元素的内容,逐单个元素匹配无法得到完整结果。
修复后的代码
import re import pandas as pd text_pos = [['5. qwe', 'LLL LLL 23', 'zzz qqq ewq (qwe ewq)', 'ewq \nqwe', 'eee wwww', 'qwewww'], ['LLL LLL 54', 'ttt qqq (eee www)', 'eeee\neee', 'aaaaa \nwww'], ['K K K K K K K K K K K K K 7 /', '111', 'zzz qqq qwe (ewq Lee)', 'qwee\neen', 'eewwww']] data = [] for block in text_pos: number_val = "" asker_val = "" # 提取number字段 for idx, s in enumerate(block): # 匹配LLL LLL开头的编号 b_match = re.search(r'LLL\s+LLL\s+(\d+)', s) if b_match: number_val = b_match.group(1).strip() break # 匹配K K K开头的编号,需要拼接下一个元素内容 a_match = re.search(r'(K\s+){13}\s*(\d.*)', s) if a_match: number_val = a_match.group(2).strip() + block[idx+1].strip() break # 提取asker字段 for s in block: c_match = re.search(r'(?:zzz|ttt)\s+(qqq.*?\))', s) if c_match: asker_val = c_match.group(1).strip() break data.append([number_val, asker_val]) df1 = pd.DataFrame(data, columns=['number', 'asker'])
执行后打印输出即可完全匹配预期结果:
# 打印number列 for v in df1['number']: print(v) # 输出: 23 54 7 /111
# 打印asker列 for v in df1['asker']: print(v) # 输出: qqq ewq (qwe ewq) qqq (eee www) qqq qwe (ewq Lee)
内容的提问来源于stack exchange,提问作者id345678
相关产品推荐
相关产品推荐

