Python+Selenium爬取纳斯达克股票:Stale Element Reference与数据重复问题
纳斯达克网站股票爬取问题解决方案
问题概述
- 翻页后数据重复:第1、2页,第3、4页提取的股票代码完全一致;页面URL无变化,通过定位页码按钮点击翻页,未使用
driver.refresh()。 - 随机触发
Stale element reference错误:在执行page_button.click()前随机出现(多在第3-5页),已添加异常捕获但无法修复上下文。
解决方案
1. 数据重复问题修复
核心原因是元素定位错误:原代码用<th>标签定位,而<th>是表头元素,每次都会重复获取表头的固定内容;同时页面动态加载后未等待DOM更新,导致获取旧页面数据。
- 修正定位:股票符号实际在
<td>标签的第一列,改用td.nasdaq-screener__cell:nth-child(1)定位 - 增加页面加载等待:确保翻页后DOM完全更新再获取源码
2. Stale元素错误修复
该错误是因为页面DOM更新后,之前缓存的元素引用失效。解决方式:
- 每次操作前重新定位元素,不复用旧的元素对象
- 使用显式等待替代隐式等待,精准控制等待条件(元素可点击、页面切换完成)
- 针对
StaleElementReferenceException增加重试逻辑
修改后的完整代码
from bs4 import BeautifulSoup as bs from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import StaleElementReferenceException, TimeoutException options = webdriver.ChromeOptions() options.add_experimental_option("detach", True) driver = webdriver.Chrome(options=options) wait = WebDriverWait(driver, 15) # 显式等待最长15秒 # NYSE 目标页面地址 url_nyse = "http://www.nasdaq.com/screening/companies-by-name.aspx?letter=0&exchange=nyse&render=download" driver.get(url_nyse) # 等待并同意隐私政策弹窗 wait.until(EC.element_to_be_clickable((By.ID, "onetrust-accept-btn-handler"))).click() # 获取总页数(可选,测试时可替换为固定数值) total_pages = wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, "pagination__page"))) max_page = int(total_pages[-1].text) # 循环爬取10页(实际可改为range(1, max_page+1)) for i in range(1, 10): try: # 等待当前页面股票数据加载完成 wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "td.nasdaq-screener__cell"))) # 获取当前页面源码并解析 html = driver.page_source soup = bs(html, "html.parser") # 精准提取第一列的股票符号 symbols = soup.select("td.nasdaq-screener__cell:nth-child(1)") for symbol in symbols: print(symbol.text.strip()) next_page_num = i + 1 print(f"准备跳转到第 {next_page_num} 页") # 重新定位下一页按钮并等待可点击 next_page_btn = wait.until(EC.element_to_be_clickable((By.XPATH, f"//button[@class='pagination__page' and text()='{next_page_num}']"))) next_page_btn.click() # 验证页面切换完成:等待目标页码变为活跃状态 wait.until(EC.presence_of_element_located((By.XPATH, f"//button[@class='pagination__page is-active' and text()='{next_page_num}']"))) print(f"已成功跳转到第 {next_page_num} 页") except StaleElementReferenceException: print(f"第 {i} 页遇到Stale元素错误,重试当前页") continue except TimeoutException: print(f"第 {i} 页等待超时,跳过") continue except Exception as e: print(f"未知错误: {str(e)}") continue driver.quit()
关键修改说明
- 修正元素定位:从表头
<th>改为数据列<td>,解决数据重复的核心问题 - 显式等待:确保页面元素加载完成后再操作,避免获取旧数据或操作未就绪的元素
- 动态定位元素:每次点击页码前重新查找元素,避免DOM更新导致的引用失效
- 页面切换验证:通过等待目标页码变为活跃状态,确认页面已完成切换
内容的提问来源于stack exchange,提问作者Salva194
相关产品推荐
相关产品推荐

