You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium 4中如何定位目标span?解决网页数据提取中断问题

解决Sun足球板块爬虫标题定位冲突问题

问题描述

我需要提取https://www.thesun.co.uk/sport/football/网站的标题、副标题和链接,但提取约30-35条数据后,部分<div class="teaser__copy-container">容器内出现两个span标签,导致当前通过./a/span定位标题的逻辑冲突,程序无法继续提取后续数据。尝试通过span的class属性定位时程序直接崩溃报错,希望得到解决办法。

原代码

import time
import pandas
from selenium import webdriver
from selenium.webdriver.chrome.service import Service as ChromeService
from webdriver_manager.chrome import ChromeDriverManager

driver = webdriver.Chrome(service=ChromeService(ChromeDriverManager().install()))

web = "https://www.thesun.co.uk/sport/football/"

driver.get(web)

containers = driver.find_elements(by="xpath", value='//div[@class="teaser__copy-container"]')

titles = []
subtitles = []
links = []

for container in containers:
    title = container.find_element(by='xpath', value='./a/span').text
    subtitle = container.find_element(by='xpath', value='./a/h3').text
    link = container.find_element(by='xpath', value='./a').get_attribute('href')

    titles.append(title)
    subtitles.append(subtitle)
    links.append(link)

my_dict = {'title': titles, 'subtitle': subtitles, 'link': links}
fd = pandas.DataFrame(my_dict)
fd.to_csv('myextract.csv')

问题分析

  • 部分容器的<a>标签下存在多个<span>节点,直接用./a/span会匹配到第一个span,但后续出现多span时,该节点可能并非标题内容
  • 按class定位崩溃是因为部分容器内的span没有目标class,find_element找不到元素时会抛出未捕获的异常,导致程序终止

解决办法及修改后代码

import time
import pandas
from selenium import webdriver
from selenium.webdriver.chrome.service import Service as ChromeService
from webdriver_manager.chrome import ChromeDriverManager
from selenium.common.exceptions import NoSuchElementException

driver = webdriver.Chrome(service=ChromeService(ChromeDriverManager().install()))
web = "https://www.thesun.co.uk/sport/football/"
driver.get(web)

# 滚动页面加载更多内容,可根据需要调整滚动次数
for _ in range(3):
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(2)

containers = driver.find_elements(by="xpath", value='//div[@class="teaser__copy-container"]')

titles = []
subtitles = []
links = []

for container in containers:
    title = ""
    try:
        # 优先定位标题对应的span(实际页面中标题span的class为teaser__headline,可根据页面结构调整)
        title_elem = container.find_element(by='xpath', value='./a/span[@class="teaser__headline"]')
        title = title_elem.text.strip()
    except NoSuchElementException:
        # 若找不到指定class的span,尝试取a标签下第一个span的文本
        try:
            title_elem = container.find_element(by='xpath', value='./a/span[1]')
            title = title_elem.text.strip()
        except NoSuchElementException:
            pass

    subtitle = ""
    try:
        subtitle_elem = container.find_element(by='xpath', value='./a/h3')
        subtitle = subtitle_elem.text.strip()
    except NoSuchElementException:
        pass

    link = ""
    try:
        link_elem = container.find_element(by='xpath', value='./a')
        link = link_elem.get_attribute('href')
    except NoSuchElementException:
        pass

    # 仅添加有有效链接的数据
    if link:
        titles.append(title)
        subtitles.append(subtitle)
        links.append(link)

my_dict = {'title': titles, 'subtitle': subtitles, 'link': links}
fd = pandas.DataFrame(my_dict)
fd.to_csv('myextract.csv')

driver.quit()

关键修改说明

  1. 异常捕获:引入NoSuchElementException捕获单个元素定位失败的情况,避免程序直接崩溃
  2. 精准定位标题:优先通过标题span的专属class(teaser__headline)定位,兼容多span场景;若找不到则退而求其次取第一个span
  3. 滚动加载:添加页面滚动逻辑,确保获取到更多初始未加载的内容
  4. 数据有效性校验:仅保留带有有效链接的数据,避免空值干扰

内容的提问来源于stack exchange,提问作者codejerry08

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 01:02:03