如何用Selenium遍历抓取网站折叠面板数据?附商店邮箱爬取问题
问题分析与解决方案
核心问题
你的代码存在以下几个关键问题导致只能获取第一个商店的邮箱:
- 陈旧元素引用:提前获取的
link_details列表,在点击第一个DETAILS后页面DOM结构发生变化(弹出详情弹窗),后续的DETAILS元素会变成无效的陈旧元素,无法再被点击。 - 未关闭详情弹窗:查看详情后没有关闭弹窗,导致后续无法定位到下一个DETAILS按钮,也无法正确加载下一个商店的详情。
- 异常处理过于粗暴:捕获所有异常就直接退出浏览器,遇到小问题就终止程序,无法继续处理后续商店。
- 数据不同步:提前单独获取所有商店名称,若后续爬取邮箱失败,会导致商店名和邮箱数量不匹配。
修复后的代码
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.keys import Keys from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import StaleElementReferenceException, NoSuchElementException import pandas as pd import time ser = Service("C:\chromedriver.exe") op = webdriver.ChromeOptions() s = webdriver.Chrome(service=ser, options=op) s.get("https://www.jimsformalwear.com/stores") search = s.find_element(by=By.ID, value='address') search.send_keys("Des Moines, IA") search.send_keys(Keys.RETURN) time.sleep(3) store_data = [] # 循环处理每个商店,每次重新定位元素避免失效 while True: try: # 获取当前页面所有商店的DETAILS按钮 link_details = WebDriverWait(s, 10).until( EC.presence_of_all_elements_located((By.LINK_TEXT, 'DETAILS')) ) # 遍历按钮,用索引遍历避免元素失效问题 for idx in range(len(link_details)): # 重新定位当前索引的DETAILS按钮,确保元素有效 detail_btn = WebDriverWait(s, 10).until( EC.element_to_be_clickable((By.LINK_TEXT, 'DETAILS')) ) # 从DETAILS按钮的父元素中获取商店名称 store_name = detail_btn.find_element(By.XPATH, './../h3').text detail_btn.click() time.sleep(2) # 获取邮箱地址,无邮箱时标记为"无邮箱" try: email = WebDriverWait(s, 10).until( EC.presence_of_element_located((By.CLASS_NAME, 'store-email.store-contact-item')) ).text except NoSuchElementException: email = "无邮箱" store_data.append({ "Store Name": store_name, "Email": email }) # 关闭详情弹窗,回到商店列表 close_btn = WebDriverWait(s, 10).until( EC.element_to_be_clickable((By.CLASS_NAME, 'close-modal')) ) close_btn.click() time.sleep(2) break except StaleElementReferenceException: # 遇到陈旧元素异常,重新定位继续循环 continue s.quit() df1 = pd.DataFrame(store_data) print(df1)
代码改进说明
- 每次重新定位元素:避免DOM变化导致的陈旧元素引用问题,循环时重新获取DETAILS按钮。
- 同步获取商店名和邮箱:从DETAILS按钮的父元素中提取商店名称,保证数据一一对应。
- 关闭详情弹窗:每次获取邮箱后点击关闭按钮,回到商店列表页面,确保后续操作正常。
- 精准异常捕获:只捕获特定的元素异常,遇到错误时跳过当前商店或重新尝试,不会直接终止程序。
- 显式等待替代固定sleep:大部分场景用WebDriverWait显式等待,提升代码稳定性和执行效率。
内容的提问来源于stack exchange,提问作者noideawhatimdoingtbh
相关产品推荐
相关产品推荐

