You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium中getAttribute("href")返回None问题求助

解决Selenium获取href返回None的问题

我看了你的代码和问题描述,能拿到标题文本但获取href返回None,大概率是你定位的元素本身不是带href的<a>标签,而是包裹<a>的父元素。咱们一步步来解决:

问题根源

你当前用find_elements_by_class_name('cnn-search__result-headline')定位的headlines_list里的元素,应该是一个容器标签(比如<div>),真正的链接<a>是它的子元素。父元素本身没有href属性,所以调用elem.get_attribute("href")自然返回None;而elem.text能拿到内容是因为Selenium会自动获取元素及其子元素的文本内容。

解决方案

有两种简单的修改方式,选一种就行:

方式1:在父元素内定位子<a>标签

保持你原来的headlines_list定位,在循环里先找到子元素<a>再拿href:

# 导入By模块(建议用新的定位方法,旧方法已被官方弃用)
from selenium.webdriver.common.by import By

# ... 你的其他代码 ...

headlines = []; links = [];
for elem in headlines_list:
    # 从容器元素里找到<a>标签
    link_element = elem.find_element(By.TAG_NAME, 'a')
    # 获取<a>的href属性
    links.append(link_element.get_attribute("href"))
    # 标题文本可以继续用elem.text,或者用link_element.text
    headlines.append(elem.text)

方式2:直接定位<a>标签

更高效的方式是直接定位到带href的<a>元素,跳过父容器:

from selenium.webdriver.common.by import By

# ... 你的其他代码 ...

# 直接通过CSS选择器定位到标题容器下的<a>标签
headline_links = main_news_container.find_elements(By.CSS_SELECTOR, '.cnn-search__result-headline a')

headlines = []; links = [];
for elem in headline_links:
    links.append(elem.get_attribute("href"))
    headlines.append(elem.text)

额外建议

  1. 尽量使用Selenium 4推荐的By类定位方法(比如By.CLASS_NAME、By.CSS_SELECTOR),你原来用的find_element_by_class_name这类旧方法已经被官方弃用了。
  2. 避免用time.sleep()这种固定等待,换成WebDriverWait显式等待,更稳定:
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 等待主容器加载完成,最多等10秒
main_news_container = WebDriverWait(browser, 10).until(
    EC.presence_of_element_located((By.CLASS_NAME, 'cnn-search__results-list'))
)

内容的提问来源于stack exchange,提问作者Mertay Dayanc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 10:27:57