使用Python .find()通过Font Awesome图标获取<a>标签href的问题排查
问题排查:通过Font Awesome图标获取对应标签的href属性
需求与HTML结构
需求:通过Font Awesome的fa-angle-right图标,使用.find()方法获取对应<a>标签的href属性,并在图标存在时打印URL及当前页面。
对应的HTML结构:
<li class="item"><a class="link" href="?page=2"><i class="fas fa-angle-right" aria-hidden="true"></i></a></li>
用户的错误代码
用户尝试的Python代码(存在错误):
def next_page(self, current_page: HTML): next_page_href = current_page.find('i',{"class": "fa-angle-right"}, first=True) next_page_html = self._get_url_content('{url}{next_page}'.format( url=self.url, next_page=next_page_href.attrs['href'] )) if current_page.url == next_page_html.url: return None return next_page_html
错误分析及修正方案
核心问题
- 元素定位错误:代码直接获取
<i>标签,但href属性在它的父节点<a>上,<i>标签本身没有href属性,直接访问会报错。 - 无异常处理:如果页面不存在目标图标,
next_page_href会是None,此时访问attrs会抛出AttributeError。 - 缺少打印逻辑:原代码未实现需求中“打印URL及当前页面”的要求。
修正后的代码
def next_page(self, current_page: HTML): # 定位fa-angle-right图标 right_icon = current_page.find('i', {"class": "fa-angle-right"}, first=True) if not right_icon: print("未找到fa-angle-right图标") return None # 获取图标对应的<a>父标签 next_link = right_icon.parent if not next_link or next_link.tag != 'a': print("图标父元素不是<a>标签,无法获取href") return None # 提取href属性 next_page_path = next_link.attrs.get('href') if not next_page_path: print("<a>标签未设置href属性") return None # 拼接完整下一页URL full_next_url = f"{self.url}{next_page_path}" next_page_html = self._get_url_content(full_next_url) # 打印当前页面和下一页URL print(f"当前页面URL: {current_page.url}") print(f"下一页URL: {full_next_url}") if current_page.url == next_page_html.url: return None return next_page_html
关键改动说明
- 先定位
<i>标签,再通过.parent获取其父节点<a>,这才是包含href属性的目标元素 - 增加多层异常判断,覆盖图标不存在、父元素非
<a>、<a>无href的场景,避免程序崩溃 - 新增打印逻辑,满足需求中输出URL的要求
- 用f-string简化URL拼接,写法更直观简洁
内容的提问来源于stack exchange,提问作者JasSum
相关产品推荐
相关产品推荐

