Selenium 4中如何定位目标span?解决网页数据提取中断问题
解决Sun足球板块爬虫标题定位冲突问题
问题描述
我需要提取https://www.thesun.co.uk/sport/football/网站的标题、副标题和链接,但提取约30-35条数据后,部分<div class="teaser__copy-container">容器内出现两个span标签,导致当前通过./a/span定位标题的逻辑冲突,程序无法继续提取后续数据。尝试通过span的class属性定位时程序直接崩溃报错,希望得到解决办法。
原代码
import time import pandas from selenium import webdriver from selenium.webdriver.chrome.service import Service as ChromeService from webdriver_manager.chrome import ChromeDriverManager driver = webdriver.Chrome(service=ChromeService(ChromeDriverManager().install())) web = "https://www.thesun.co.uk/sport/football/" driver.get(web) containers = driver.find_elements(by="xpath", value='//div[@class="teaser__copy-container"]') titles = [] subtitles = [] links = [] for container in containers: title = container.find_element(by='xpath', value='./a/span').text subtitle = container.find_element(by='xpath', value='./a/h3').text link = container.find_element(by='xpath', value='./a').get_attribute('href') titles.append(title) subtitles.append(subtitle) links.append(link) my_dict = {'title': titles, 'subtitle': subtitles, 'link': links} fd = pandas.DataFrame(my_dict) fd.to_csv('myextract.csv')
问题分析
- 部分容器的
<a>标签下存在多个<span>节点,直接用./a/span会匹配到第一个span,但后续出现多span时,该节点可能并非标题内容 - 按class定位崩溃是因为部分容器内的span没有目标class,
find_element找不到元素时会抛出未捕获的异常,导致程序终止
解决办法及修改后代码
import time import pandas from selenium import webdriver from selenium.webdriver.chrome.service import Service as ChromeService from webdriver_manager.chrome import ChromeDriverManager from selenium.common.exceptions import NoSuchElementException driver = webdriver.Chrome(service=ChromeService(ChromeDriverManager().install())) web = "https://www.thesun.co.uk/sport/football/" driver.get(web) # 滚动页面加载更多内容,可根据需要调整滚动次数 for _ in range(3): driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(2) containers = driver.find_elements(by="xpath", value='//div[@class="teaser__copy-container"]') titles = [] subtitles = [] links = [] for container in containers: title = "" try: # 优先定位标题对应的span(实际页面中标题span的class为teaser__headline,可根据页面结构调整) title_elem = container.find_element(by='xpath', value='./a/span[@class="teaser__headline"]') title = title_elem.text.strip() except NoSuchElementException: # 若找不到指定class的span,尝试取a标签下第一个span的文本 try: title_elem = container.find_element(by='xpath', value='./a/span[1]') title = title_elem.text.strip() except NoSuchElementException: pass subtitle = "" try: subtitle_elem = container.find_element(by='xpath', value='./a/h3') subtitle = subtitle_elem.text.strip() except NoSuchElementException: pass link = "" try: link_elem = container.find_element(by='xpath', value='./a') link = link_elem.get_attribute('href') except NoSuchElementException: pass # 仅添加有有效链接的数据 if link: titles.append(title) subtitles.append(subtitle) links.append(link) my_dict = {'title': titles, 'subtitle': subtitles, 'link': links} fd = pandas.DataFrame(my_dict) fd.to_csv('myextract.csv') driver.quit()
关键修改说明
- 异常捕获:引入
NoSuchElementException捕获单个元素定位失败的情况,避免程序直接崩溃 - 精准定位标题:优先通过标题span的专属class(
teaser__headline)定位,兼容多span场景;若找不到则退而求其次取第一个span - 滚动加载:添加页面滚动逻辑,确保获取到更多初始未加载的内容
- 数据有效性校验:仅保留带有有效链接的数据,避免空值干扰
内容的提问来源于stack exchange,提问作者codejerry08
相关产品推荐
相关产品推荐

