Python正则匹配HTML链接时出现AttributeError: 'NoneType' object has no attribute 'group'错误
Python正则匹配HTML链接时出现AttributeError: 'NoneType' object has no attribute 'group'错误
这个问题我碰到过,其实核心原因是你的正则在Python里没找到匹配内容,返回了None,直接调用group()肯定就触发这个报错啦。咱们一步步拆解问题和解决办法:
为什么Notepad++能匹配但Python不行?
Notepad++的正则默认是多行模式,也就是^会匹配每一行的开头;但Python里的re.search和re.match默认是把整个字符串当成单一行处理,^只会匹配整个字符串的开头。你的目标<a>标签前面还有一行<div>内容,所以默认模式下^\s*根本匹配不到目标行。
解决办法
1. 开启多行模式
给正则加上re.MULTILINE(或简写re.M)标志,让^匹配每一行的开头:
import re # 加上re.MULTILINE标志 match_result = re.search(r'^\s*<a href="(.*?)" title="View', new_file_content, re.MULTILINE) if match_result: b_content = match_result.group(1) # 执行后续替换操作 old_file_content = re.sub(', in <a href="(.*?)" title="Vezi', f', in <a href="{b_content}" title="Vezi', old_file_content) else: print("没找到符合条件的链接哦")
2. 去掉行开头限制(如果场景允许)
如果你的文件里只有这一个符合title以View开头的<a>标签,也可以直接去掉^\s*,让re.search扫描整个字符串找匹配:
match_result = re.search(r'<a href="(.*?)" title="View', new_file_content) if match_result: b_content = match_result.group(1)
3. 更靠谱的方案:用HTML解析库代替正则
正则处理HTML很容易因为标签格式变化(比如空格、属性顺序变了)失效,推荐用专门的HTML解析库,比如BeautifulSoup,代码更稳定:
from bs4 import BeautifulSoup soup = BeautifulSoup(new_file_content, 'html.parser') # 找到title属性以"View"开头的<a>标签 target_link = soup.find('a', title=lambda t: t and t.startswith('View')) if target_link: b_content = target_link['href'] # 执行替换 old_file_content = re.sub(', in <a href="(.*?)" title="Vezi', f', in <a href="{b_content}" title="Vezi', old_file_content)
备注:内容来源于stack exchange,提问作者Hellena Crainicu
相关产品推荐
相关产品推荐

