You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则Match匹配失败但regex101正常,求问题排查

为什么你的Python正则匹配失败?

嘿,我一眼就看出问题所在了——你在Python里用了re.match(),但这个方法的行为和你用的测试工具不一样!

核心问题

re.match()只从字符串的起始位置开始匹配,而你的HTML字符串开头是<!DOCTYPE html>,根本不是<title>标签,所以匹配直接失败。但测试工具默认是在整个字符串里搜索匹配项,这相当于Python中的re.search()方法,所以能正常找到结果。

快速修复

把代码里的match = title_pattern.match(html)替换成match = title_pattern.search(html)就行。

修改后的完整代码:

import re

def get_title_and_content(html):
    html = """<!DOCTYPE html> <html> <head> <title>Change delivery date with Deliv</title> </head> <body> <div class="gkms web">The delivery date can be changed up until the package is assigned to a driver.</div> </body> </html> """
    title_pattern = re.compile(r'<title>(.*?)</title>(.*)')
    match = title_pattern.search(html)
    if match:
        print('successfully extract title and answer')
        return match.groups()[0].strip(), match.groups()[1].strip()
    else:
        print('unable to extract title or answer')

额外提醒

  • 区分match()和search():match()锚定字符串开头,search()扫描整个字符串找第一个匹配项;
  • 解析HTML更推荐用BeautifulSoup这类专门的库,正则处理HTML很容易因为标签格式变化(比如带属性、换行)失效。

内容的提问来源于stack exchange,提问作者Yang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:24:17