You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium获取href循环中断问题求助

问题分析

你遇到的NoSuchElementException是因为不是所有teaser__copy-container容器内部都包含<a>标签,当循环到没有链接的容器时,find_element找不到元素就会直接抛出错误中断循环。

解决方案

可以通过两种方式解决:

  • 用try-except捕获异常,遇到无链接的容器时跳过或标记为None
  • 先筛选出包含<a>标签的容器再循环

另外,建议添加显式等待确保页面元素完全加载,避免因元素未加载导致的查找失败。

修改后的代码

import pandas as pd
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

website = 'https://www.thesun.co.uk/sport/football/'
path = 'D:\\Programming\\Automate with Python\\Automating\\chromedriver_win32'
chrome_options = Options()
chrome_options.add_experimental_option('detach', True)

service = Service(executable_path=path)
browser = webdriver.Chrome(options=chrome_options, service=service)
browser.get(website)

# 显式等待容器加载完成,最多等待10秒
wait = WebDriverWait(browser, 10)
containers = wait.until(EC.presence_of_all_elements_located((By.XPATH, '//div[@class="teaser__copy-container"]')))

titles = []
sub_titles = []
links = []

for container in containers:
    try:
        title = container.find_element(By.CSS_SELECTOR, 'span').text
        sub_title = container.find_element(By.CSS_SELECTOR, 'h3').text
        # 尝试查找链接,找不到则抛出异常进入except
        link = container.find_element(By.XPATH, './/a').get_attribute("href")
        
        titles.append(title)
        sub_titles.append(sub_title)
        links.append(link)
        print(link)
    except Exception as e:
        # 打印异常信息(可选),跳过当前无链接的容器
        print(f"跳过无链接的容器: {str(e)}")
        continue

df_headlines = pd.DataFrame({'title': titles, 'sub-title': sub_titles, 'links': links})
df_headlines.to_csv('headline.csv')
browser.quit()

关键修改点

  • 添加显式等待:用WebDriverWait确保所有容器加载完成,避免因页面未完全渲染导致的元素遗漏
  • 加入try-except块:捕获元素查找异常,遇到无链接的容器时自动跳过,保证循环正常执行
  • 替换get_attribute("textContent")为.text:Selenium的.text属性更简洁,直接获取元素可见文本
  • 最后调用browser.quit():关闭浏览器释放资源

另一种优化思路:直接筛选含链接的容器

如果只想爬取有链接的内容,可以直接修改容器的定位XPath,只选择内部包含<a>标签的容器:

containers = wait.until(EC.presence_of_all_elements_located((By.XPATH, '//div[@class="teaser__copy-container" and .//a]')))

这样循环时每个容器都必然有<a>标签,无需额外异常处理。

内容的提问来源于stack exchange,提问作者Sam Schmierer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 00:47:09