Python使用re正则从HTML字符串列表提取指定内容代码失效求助
问题原因
- 正则表达式中的
.*默认采用贪婪匹配规则,会匹配尽可能长的字符范围:会从字符串第一个>开始,一直匹配到整个字符串最后一个<才停止,最终拿到的内容不符合预期。 - 没有做匹配结果判空处理,若某条字符串不符合匹配规则,直接调用
.group(1)会抛出异常。
修复后可运行代码
import re start = re.escape(">") end = re.escape("<") stringlist =['<div class="ant-space-item"><a href="/holdings-of-1">myinformation_1</a></div>', '<div class="ant-space-item"><a href="/holdings-of-2avbf">myinformation_2</a></div>'] for i in stringlist : # 把.*改为.*?开启非贪婪匹配,匹配到最近的<就停止 match_res = re.search('%s(.*?)%s' % (start, end), i) if match_res: result = match_res.group(1) # 过滤掉标签之间为空的情况 if result: print(result)
运行输出
myinformation_1 myinformation_2
优化方案
如果固定提取a标签内的文本,可以写更精准的正则减少误匹配:
import re stringlist =['<div class="ant-space-item"><a href="/holdings-of-1">myinformation_1</a></div>', '<div class="ant-space-item"><a href="/holdings-of-2avbf">myinformation_2</a></div>'] for i in stringlist: match_res = re.search(r'<a[^>]*>(.*?)</a>', i) if match_res: print(match_res.group(1))
内容的提问来源于stack exchange,提问作者Hary2
相关产品推荐
相关产品推荐

