You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用re正则从HTML字符串列表提取指定内容代码失效求助

问题原因

  • 正则表达式中的.*默认采用贪婪匹配规则,会匹配尽可能长的字符范围:会从字符串第一个>开始,一直匹配到整个字符串最后一个<才停止,最终拿到的内容不符合预期。
  • 没有做匹配结果判空处理,若某条字符串不符合匹配规则,直接调用.group(1)会抛出异常。

修复后可运行代码

import re
start = re.escape(">")
end   = re.escape("<")
stringlist =['<div class="ant-space-item"><a href="/holdings-of-1">myinformation_1</a></div>', 
    '<div class="ant-space-item"><a href="/holdings-of-2avbf">myinformation_2</a></div>']
for i in stringlist :
    # 把.*改为.*?开启非贪婪匹配,匹配到最近的<就停止
    match_res = re.search('%s(.*?)%s' % (start, end), i)
    if match_res:
        result = match_res.group(1)
        # 过滤掉标签之间为空的情况
        if result:
            print(result)

运行输出

myinformation_1
myinformation_2

优化方案

如果固定提取a标签内的文本,可以写更精准的正则减少误匹配:

import re
stringlist =['<div class="ant-space-item"><a href="/holdings-of-1">myinformation_1</a></div>', 
    '<div class="ant-space-item"><a href="/holdings-of-2avbf">myinformation_2</a></div>']
for i in stringlist:
    match_res = re.search(r'<a[^>]*>(.*?)</a>', i)
    if match_res:
        print(match_res.group(1))

内容的提问来源于stack exchange,提问作者Hary2

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 04:36:04