You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python-BeautifulSoup抓取页面中的多个href链接?

使用Python BeautifulSoup抓取指定链接的正确方法

你的代码存在两处关键错误:

  1. 错误使用findAll('div', {'a class': 'xxx'}),这个写法是查找带有a class属性的div元素,而非div下的目标a标签;
  2. make_list是div元素的列表,无法直接调用findAll('href')——href是a标签的属性,不是独立标签。

正确实现方式

方法一:直接定位目标a标签

直接筛选所有符合class的a标签,再提取href属性:

# 查找所有目标a标签
link_items = base_soup.find_all('a', class_='link--muted no--text--decoration result-item')

# 提取有效href属性(过滤空值)
links = [item.get('href') for item in link_items if item.get('href')]

print(links)

方法二:先定位父div再找a标签(更精准)

如果页面中存在其他相同class的a标签,可先定位到包含目标链接的父div,再查找内部的a标签:

# 定位父div元素
target_div = base_soup.find('div', class_='cBox-body cBox-body--eyeCatcher')

if target_div:
    # 在父div内查找目标a标签
    link_items = target_div.find_all('a', class_='link--muted no--text--decoration result-item')
    # 提取href属性
    links = [item.get('href') for item in link_items if item.get('href')]
    print(links)

关键说明

  • 使用class_作为指定class的参数(避免与Python内置class关键字冲突);
  • 用tag.get('href')获取属性值,比tag['href']更安全——当标签无href属性时,get()会返回None而不是抛出异常;
  • 列表推导式可快速生成有效链接列表。

内容的提问来源于stack exchange,提问作者test kullanıcısı

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 16:20:35