爬取网页href时出现AttributeError: 'NoneType'无attrs属性错误求解
错误原因与修复方案
这个错误的核心是:你遍历的部分section元素里没有找到<a>标签,item.find('a', first=True)返回了None,而None没有attrs属性,所以触发了AttributeError。
修复方法
在获取href前先判断是否存在<a>标签,或者用安全的方式访问属性,下面是两种可行的修改方案:
方案1:添加存在性判断
from requests_html import HTMLSession s = HTMLSession() url = 'https://lakesshoweringspaces.com/catalogue-product-filter/page/1' r = s.get(url) products = r.html.find('article.contentwrapper section') for item in products: # 先获取a标签对象,判断是否存在 link = item.find('a', first=True) if link: # 如果link不是None,再获取href print(link.attrs['href'])
方案2:使用字典get方法(更简洁)
如果想避免显式判断,也可以用get方法安全获取属性,不存在时返回默认值(比如None):
from requests_html import HTMLSession s = HTMLSession() url = 'https://lakesshoweringspaces.com/catalogue-product-filter/page/1' r = s.get(url) products = r.html.find('article.contentwrapper section') for item in products: link = item.find('a', first=True) # 用get获取href,不存在时返回None,不会报错 print(link.attrs.get('href') if link else None)
额外建议
你可以先打印所有products的内容,看看哪些section里没有<a>标签,这样能更精准地调整选择器(比如使用更具体的选择器,如article.contentwrapper section.product-item,过滤掉不含链接的无效元素)。
内容的提问来源于stack exchange,提问作者HRol
相关产品推荐
相关产品推荐

