使用Selenium抓取参展商信息时点击跳转报错求助
我需要从https://asiatechxsg.com/exhibitors/抓取所有参展商名称及相关信息并导出为CSV文件。目前用Selenium写的代码能识别参展商链接,但点击进入详情页提取信息时报错,错误栈如下:
INFO:root:Found 860 exhibitor links. INFO:root:Clicked on exhibitor link 1/860 ERROR:root:Failed to process exhibitor link: Message: Stacktrace: GetHandleVerifier [0x00007FF65B8C1522+60802] (No symbol) [0x00007FF65B83AC22] (No symbol) [0x00007FF65B6F7CE4] (No symbol) [0x00007FF65B746D4D] (No symbol) [0x00007FF65B746E1C] (No symbol) [0x00007FF65B78CE37] (No symbol) [0x00007FF65B76ABBF] (No symbol) [0x00007FF65B78A224] (No symbol) [0x00007FF65B76A923] (No symbol) [0x00007FF65B738FEC] (No symbol) [0x00007FF65B739C21] GetHandleVerifier [0x00007FF65BBC41BD+3217949] GetHandleVerifier [0x00007FF65BC06157+3488183]
我是Selenium新手,不清楚问题所在,请求帮助解决错误,完成所有参展商信息的抓取。现有代码如下:
exhibitor_names = [] exhibitor_info = [] exhibitor_web = [] exhibitor_linkedin = [] exhibitor_twitter = [] exhibitor_email = [] exhibitor_contact = [] exhibitor_booth = [] import time import logging import pandas as pd from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.service import Service from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Configure logging logging.basicConfig(level=logging.INFO) logger = logging.getLogger() # Initialize WebDriver and Options driver = webdriver.Chrome() driver.maximize_window() driver.get("https://asiatechxsg.com/exhibitors/") wait = WebDriverWait(driver, 10) wait.until(EC.frame_to_be_available_and_switch_to_it((By.CLASS_NAME, "IframeModule_iframe__JCvXg"))) driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") match=False while(match==False): driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") last_element = driver.find_elements(By.XPATH, "//img[@alt='Singapore Centre for Social Enterprise, raiSE Ltd']") if (len(last_element) >= 1): match=True # Extract all exhibitor links exhibitor_links = driver.find_elements(By.XPATH, "//a[contains(@href, '/widget/event/asia-tech-x-singapore-2024/exhibitor/')]") logger.info(f"Found {len(exhibitor_links)} exhibitor links.") for index, link in enumerate(exhibitor_links): try: # Click on each exhibitor link link.click() logger.info(f"Clicked on exhibitor link {index + 1}/{len(exhibitor_links)}") # Wait for the div containing the paragraphs to be visible info_div = wait.until(EC.visibility_of_element_located((By.XPATH, "//div[@class='sc-dd6f9f7c-0 hIHCKf']"))) # Extract all the paragraph texts within the div paragraphs = info_div.find_elements(By.TAG_NAME, "p") info_text = '\n'.join([p.text for p in paragraphs]) # Extract the exhibitor name name = driver.find_element(By.XPATH, "//span[@class='sc-a13c392f-0 sc-85df61db-4 feepZZ eErKol']").text # Append the extracted information to the lists exhibitor_names.append(name) exhibitor_info.append(info_text) logger.info(f"Extracted information for {name}.") # Navigate back to the main exhibitors page driver.back() logger.info("Navigated back to the main exhibitors page.") except Exception as e: logger.error("Failed to process exhibitor link: %s", e) continue # Close the WebDriver driver.quit() # Create a DataFrame exhibitors_df = pd.DataFrame({ 'Exhibitor Name': exhibitor_names, 'Description': exhibitor_info, }) # Print the extracted information print(exhibitors_df)
问题根源
- 元素引用失效:点击链接后页面跳转,原
link元素的DOM引用失效,触发StaleElementReferenceException。 - 动态类名不稳定:代码中使用的
sc-dd6f9f7c-0 hIHCKf这类随机生成的类名,页面刷新后会变化,导致元素定位失败。 - iframe上下文未重置:详情页同样在iframe内,跳转后未重新切换iframe上下文,导致无法找到元素。
- 页面加载等待不足:点击链接后直接等待元素,未给页面足够加载时间,或未判断新页面的iframe是否加载完成。
修复步骤
1. 改用URL跳转替代点击元素
先提取所有链接的href属性,再逐个用driver.get()打开,避免元素引用失效问题:
# 替换原链接提取逻辑 exhibitor_links = [link.get_attribute('href') for link in driver.find_elements(By.XPATH, "//a[contains(@href, '/widget/event/asia-tech-x-singapore-2024/exhibitor/')]")] logger.info(f"Found {len(exhibitor_links)} exhibitor links.") # 遍历链接列表 for index, url in enumerate(exhibitor_links): try: driver.get(url) logger.info(f"Processing exhibitor {index + 1}/{len(exhibitor_links)}")
2. 替换动态类名为稳定定位方式
放弃随机类名,改用结构或文本特征定位:
- 参展商名称:用
<h1>标签或父容器相对路径(如//div[contains(@class, 'exhibitor-header')]/h1) - 详情内容:用
//div[contains(@class, 'exhibitor-description')]或直接定位所有<p>标签的父容器
3. 重置iframe上下文
进入详情页后,重新等待并切换到对应的iframe:
# 打开详情页后添加 wait.until(EC.frame_to_be_available_and_switch_to_it((By.CLASS_NAME, "IframeModule_iframe__JCvXg")))
4. 完善加载等待逻辑
在打开详情页后,增加页面加载等待,比如判断标题是否加载完成:
wait.until(EC.title_contains("Exhibitor"))
5. 补充字段提取与异常处理
为避免数据错位,异常时填充空值,同时补充官网、LinkedIn等字段的稳定提取逻辑:
# 提取官网链接示例 try: website = driver.find_element(By.XPATH, "//a[contains(@href, 'http') and not(contains(@href, 'linkedin'))]").get_attribute('href') except: website = None exhibitor_web.append(website)
完整修复后代码示例
exhibitor_names = [] exhibitor_info = [] exhibitor_web = [] exhibitor_linkedin = [] exhibitor_twitter = [] exhibitor_email = [] exhibitor_contact = [] exhibitor_booth = [] import time import logging import pandas as pd from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Configure logging logging.basicConfig(level=logging.INFO) logger = logging.getLogger() # Initialize WebDriver and Options chrome_options = Options() chrome_options.add_argument("--start-maximized") driver = webdriver.Chrome(options=chrome_options) wait = WebDriverWait(driver, 15) # 主页面加载与iframe切换 driver.get("https://asiatechxsg.com/exhibitors/") wait.until(EC.frame_to_be_available_and_switch_to_it((By.CLASS_NAME, "IframeModule_iframe__JCvXg"))) # 滚动加载所有参展商 while True: prev_height = driver.execute_script("return document.body.scrollHeight") driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(2) new_height = driver.execute_script("return document.body.scrollHeight") last_element = driver.find_elements(By.XPATH, "//img[@alt='Singapore Centre for Social Enterprise, raiSE Ltd']") if len(last_element) >= 1 or new_height == prev_height: break # 提取所有参展商链接URL exhibitor_links = [link.get_attribute('href') for link in driver.find_elements(By.XPATH, "//a[contains(@href, '/widget/event/asia-tech-x-singapore-2024/exhibitor/')]")] logger.info(f"Found {len(exhibitor_links)} exhibitor links.") # 遍历每个参展商页面 for index, url in enumerate(exhibitor_links): try: driver.get(url) logger.info(f"Processing exhibitor {index + 1}/{len(exhibitor_links)}") # 切换到详情页的iframe wait.until(EC.frame_to_be_available_and_switch_to_it((By.CLASS_NAME, "IframeModule_iframe__JCvXg"))) # 提取参展商名称 name = wait.until(EC.visibility_of_element_located((By.TAG_NAME, "h1"))).text exhibitor_names.append(name) # 提取详情描述 info_div = wait.until(EC.visibility_of_element_located((By.XPATH, "//div[contains(@class, 'sc-dd6f9f7c-0')]"))) paragraphs = info_div.find_elements(By.TAG_NAME, "p") info_text = '\n'.join([p.text.strip() for p in paragraphs if p.text.strip()]) exhibitor_info.append(info_text) # 提取官网链接 try: web = driver.find_element(By.XPATH, "//a[contains(@href, 'http') and not(contains(@href, 'linkedin')) and not(contains(@href, 'twitter'))]").get_attribute('href') except: web = None exhibitor_web.append(web) # 提取LinkedIn链接 try: linkedin = driver.find_element(By.XPATH, "//a[contains(@href, 'linkedin')]").get_attribute('href') except: linkedin = None exhibitor_linkedin.append(linkedin) # 提取展位号 try: booth = driver.find_element(By.XPATH, "//div[contains(text(), 'Booth')]/following-sibling::div").text except: booth = None exhibitor_booth.append(booth) logger.info(f"Successfully extracted data for {name}") except Exception as e: logger.error(f"Failed to process exhibitor {index + 1}: {str(e)}") # 填充空值避免数据错位 exhibitor_names.append(None) exhibitor_info.append(None) exhibitor_web.append(None) exhibitor_linkedin.append(None) exhibitor_booth.append(None) continue # 关闭浏览器 driver.quit() # 创建DataFrame并导出CSV exhibitors_df = pd.DataFrame({ 'Exhibitor Name': exhibitor_names, 'Description': exhibitor_info, 'Website': exhibitor_web, 'LinkedIn': exhibitor_linkedin, 'Booth Number': exhibitor_booth }) # 过滤空值记录 exhibitors_df = exhibitors_df.dropna(subset=['Exhibitor Name']) # 导出CSV exhibitors_df.to_csv('asiatechx_exhibitors.csv', index=False, encoding='utf-8-sig') logger.info(f"Data exported to asiatechx_exhibitors.csv, total records: {len(exhibitors_df)}")
注意事项
- 动态类名可能随时变化,每次运行前需用浏览器开发者工具确认元素选择器
- 可适当调整等待超时时间或增加
time.sleep(),适配网络加载速度 - 建议添加请求间隔(如每次页面跳转后等待1-2秒),避免触发网站反爬限制
内容的提问来源于stack exchange,提问作者user22279494

