You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python+Selenium爬取纳斯达克股票:Stale Element Reference与数据重复问题

纳斯达克网站股票爬取问题解决方案

问题概述

  • 翻页后数据重复:第1、2页,第3、4页提取的股票代码完全一致;页面URL无变化,通过定位页码按钮点击翻页,未使用driver.refresh()。
  • 随机触发Stale element reference错误:在执行page_button.click()前随机出现(多在第3-5页),已添加异常捕获但无法修复上下文。

解决方案

1. 数据重复问题修复

核心原因是元素定位错误:原代码用<th>标签定位,而<th>是表头元素,每次都会重复获取表头的固定内容;同时页面动态加载后未等待DOM更新,导致获取旧页面数据。

  • 修正定位:股票符号实际在<td>标签的第一列,改用td.nasdaq-screener__cell:nth-child(1)定位
  • 增加页面加载等待:确保翻页后DOM完全更新再获取源码

2. Stale元素错误修复

该错误是因为页面DOM更新后,之前缓存的元素引用失效。解决方式:

  • 每次操作前重新定位元素,不复用旧的元素对象
  • 使用显式等待替代隐式等待,精准控制等待条件(元素可点击、页面切换完成)
  • 针对StaleElementReferenceException增加重试逻辑

修改后的完整代码

from bs4 import BeautifulSoup as bs
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import StaleElementReferenceException, TimeoutException

options = webdriver.ChromeOptions()
options.add_experimental_option("detach", True)

driver = webdriver.Chrome(options=options)
wait = WebDriverWait(driver, 15)  # 显式等待最长15秒

# NYSE 目标页面地址
url_nyse = "http://www.nasdaq.com/screening/companies-by-name.aspx?letter=0&exchange=nyse&render=download"

driver.get(url_nyse)
# 等待并同意隐私政策弹窗
wait.until(EC.element_to_be_clickable((By.ID, "onetrust-accept-btn-handler"))).click()

# 获取总页数(可选,测试时可替换为固定数值)
total_pages = wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, "pagination__page")))
max_page = int(total_pages[-1].text)

# 循环爬取10页(实际可改为range(1, max_page+1))
for i in range(1, 10):
    try:
        # 等待当前页面股票数据加载完成
        wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "td.nasdaq-screener__cell")))
        
        # 获取当前页面源码并解析
        html = driver.page_source
        soup = bs(html, "html.parser")
        
        # 精准提取第一列的股票符号
        symbols = soup.select("td.nasdaq-screener__cell:nth-child(1)")
        for symbol in symbols:
            print(symbol.text.strip())
        
        next_page_num = i + 1
        print(f"准备跳转到第 {next_page_num} 页")
        
        # 重新定位下一页按钮并等待可点击
        next_page_btn = wait.until(EC.element_to_be_clickable((By.XPATH, f"//button[@class='pagination__page' and text()='{next_page_num}']")))
        next_page_btn.click()
        
        # 验证页面切换完成:等待目标页码变为活跃状态
        wait.until(EC.presence_of_element_located((By.XPATH, f"//button[@class='pagination__page is-active' and text()='{next_page_num}']")))
        print(f"已成功跳转到第 {next_page_num} 页")

    except StaleElementReferenceException:
        print(f"第 {i} 页遇到Stale元素错误,重试当前页")
        continue
    except TimeoutException:
        print(f"第 {i} 页等待超时,跳过")
        continue
    except Exception as e:
        print(f"未知错误: {str(e)}")
        continue

driver.quit()

关键修改说明

  • 修正元素定位:从表头<th>改为数据列<td>,解决数据重复的核心问题
  • 显式等待:确保页面元素加载完成后再操作,避免获取旧数据或操作未就绪的元素
  • 动态定位元素:每次点击页码前重新查找元素,避免DOM更新导致的引用失效
  • 页面切换验证:通过等待目标页码变为活跃状态,确认页面已完成切换

内容的提问来源于stack exchange,提问作者Salva194

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 07:15:21