You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取ADB网站时提取a标签href返回None的技术求助

解决提取a标签href返回None的问题

我看了你的代码和输出的div_tags内容,马上就找到问题根源了——页面里的div.item-title是成对出现的:一半的div里包裹着带href的a标签,另一半只是纯文本的重复标题。所以当你循环遍历所有div时,遇到那些没有a标签的div,tags.find('a')会返回None,这时候调用get('href')自然拿不到有效值,甚至会触发AttributeError。

给你两种解决方案,都能完美解决这个问题:

方案1:过滤无a标签的div

在循环里先判断a标签是否存在,只处理有a标签的div:

from selenium import webdriver
import time
from bs4 import BeautifulSoup

driver = webdriver.Chrome(r"E:\chromedriver_win32\chromedriver.exe")
url= "https://www.adb.org/projects/tenders/sector/information-and-communication-technology-1066"
driver.get(url)
content = driver.page_source.encode('utf-8').strip()
soup = BeautifulSoup(content,"html.parser")
div_tags = soup.findAll("div",{"class":"item-title"})

for tags in div_tags:
    a_tag = tags.find('a')
    # 跳过没有a标签的div
    if not a_tag:
        continue
    link = a_tag.get('href')
    # 把相对路径拼接成完整URL(可选,但更实用)
    full_link = "https://www.adb.org" + link
    print(full_link)

方案2:直接定位目标a标签

更高效的方式是直接用CSS选择器定位所有div.item-title下的a标签,跳过那些无a的div:

from selenium import webdriver
import time
from bs4 import BeautifulSoup

driver = webdriver.Chrome(r"E:\chromedriver_win32\chromedriver.exe")
url= "https://www.adb.org/projects/tenders/sector/information-and-communication-technology-1066"
driver.get(url)
content = driver.page_source.encode('utf-8').strip()
soup = BeautifulSoup(content,"html.parser")

# 直接抓取所有符合条件的a标签
a_tags = soup.select("div.item-title a")
for a_tag in a_tags:
    link = a_tag.get('href')
    full_link = "https://www.adb.org" + link
    print(full_link)

补充说明:页面出现重复的div.item-title大概率是前端为了响应式布局做的处理(比如在移动端显示纯文本标题,桌面端显示可点击的链接),所以我们只需要抓取带a标签的那部分即可。

内容的提问来源于stack exchange,提问作者jjnair

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 20:22:55