You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium抓取参展商信息时点击跳转报错求助

问题:Selenium抓取参展商详情页报错

我需要从https://asiatechxsg.com/exhibitors/抓取所有参展商名称及相关信息并导出为CSV文件。目前用Selenium写的代码能识别参展商链接,但点击进入详情页提取信息时报错,错误栈如下:

INFO:root:Found 860 exhibitor links. INFO:root:Clicked on exhibitor link 1/860 ERROR:root:Failed to process exhibitor link: Message: Stacktrace: GetHandleVerifier [0x00007FF65B8C1522+60802] (No symbol) [0x00007FF65B83AC22] (No symbol) [0x00007FF65B6F7CE4] (No symbol) [0x00007FF65B746D4D] (No symbol) [0x00007FF65B746E1C] (No symbol) [0x00007FF65B78CE37] (No symbol) [0x00007FF65B76ABBF] (No symbol) [0x00007FF65B78A224] (No symbol) [0x00007FF65B76A923] (No symbol) [0x00007FF65B738FEC] (No symbol) [0x00007FF65B739C21] GetHandleVerifier [0x00007FF65BBC41BD+3217949] GetHandleVerifier [0x00007FF65BC06157+3488183]

我是Selenium新手,不清楚问题所在,请求帮助解决错误,完成所有参展商信息的抓取。现有代码如下:

exhibitor_names = []
exhibitor_info = []
exhibitor_web = []
exhibitor_linkedin = []
exhibitor_twitter = []
exhibitor_email = []
exhibitor_contact = []
exhibitor_booth = []

import time
import logging
import pandas as pd
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# Configure logging
logging.basicConfig(level=logging.INFO)
logger = logging.getLogger()

# Initialize WebDriver and Options
driver = webdriver.Chrome()
driver.maximize_window()
driver.get("https://asiatechxsg.com/exhibitors/")

wait = WebDriverWait(driver, 10)
wait.until(EC.frame_to_be_available_and_switch_to_it((By.CLASS_NAME, "IframeModule_iframe__JCvXg")))
driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
match=False
while(match==False):
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    last_element = driver.find_elements(By.XPATH, "//img[@alt='Singapore Centre for Social Enterprise, raiSE Ltd']")
    if (len(last_element) >= 1):
        match=True

# Extract all exhibitor links
exhibitor_links = driver.find_elements(By.XPATH, "//a[contains(@href, '/widget/event/asia-tech-x-singapore-2024/exhibitor/')]")
logger.info(f"Found {len(exhibitor_links)} exhibitor links.")

for index, link in enumerate(exhibitor_links):
    try:
        # Click on each exhibitor link
        link.click()
        logger.info(f"Clicked on exhibitor link {index + 1}/{len(exhibitor_links)}")

        # Wait for the div containing the paragraphs to be visible
        info_div = wait.until(EC.visibility_of_element_located((By.XPATH, "//div[@class='sc-dd6f9f7c-0 hIHCKf']")))
        
        # Extract all the paragraph texts within the div
        paragraphs = info_div.find_elements(By.TAG_NAME, "p")
        info_text = '\n'.join([p.text for p in paragraphs])
        
        # Extract the exhibitor name
        name = driver.find_element(By.XPATH, "//span[@class='sc-a13c392f-0 sc-85df61db-4 feepZZ eErKol']").text

        # Append the extracted information to the lists
        exhibitor_names.append(name)
        exhibitor_info.append(info_text)
        logger.info(f"Extracted information for {name}.")

        # Navigate back to the main exhibitors page
        driver.back()
        logger.info("Navigated back to the main exhibitors page.")
    except Exception as e:
        logger.error("Failed to process exhibitor link: %s", e)
        continue

# Close the WebDriver
driver.quit()

# Create a DataFrame
exhibitors_df = pd.DataFrame({
    'Exhibitor Name': exhibitor_names,
    'Description': exhibitor_info,
})

# Print the extracted information
print(exhibitors_df)

解决方案

问题根源

  1. 元素引用失效:点击链接后页面跳转,原link元素的DOM引用失效,触发StaleElementReferenceException。
  2. 动态类名不稳定:代码中使用的sc-dd6f9f7c-0 hIHCKf这类随机生成的类名,页面刷新后会变化,导致元素定位失败。
  3. iframe上下文未重置:详情页同样在iframe内,跳转后未重新切换iframe上下文,导致无法找到元素。
  4. 页面加载等待不足:点击链接后直接等待元素,未给页面足够加载时间,或未判断新页面的iframe是否加载完成。

修复步骤

1. 改用URL跳转替代点击元素

先提取所有链接的href属性,再逐个用driver.get()打开,避免元素引用失效问题:

# 替换原链接提取逻辑
exhibitor_links = [link.get_attribute('href') for link in driver.find_elements(By.XPATH, "//a[contains(@href, '/widget/event/asia-tech-x-singapore-2024/exhibitor/')]")]
logger.info(f"Found {len(exhibitor_links)} exhibitor links.")

# 遍历链接列表
for index, url in enumerate(exhibitor_links):
    try:
        driver.get(url)
        logger.info(f"Processing exhibitor {index + 1}/{len(exhibitor_links)}")

2. 替换动态类名为稳定定位方式

放弃随机类名,改用结构或文本特征定位:

  • 参展商名称:用<h1>标签或父容器相对路径(如//div[contains(@class, 'exhibitor-header')]/h1)
  • 详情内容:用//div[contains(@class, 'exhibitor-description')]或直接定位所有<p>标签的父容器

3. 重置iframe上下文

进入详情页后,重新等待并切换到对应的iframe:

# 打开详情页后添加
wait.until(EC.frame_to_be_available_and_switch_to_it((By.CLASS_NAME, "IframeModule_iframe__JCvXg")))

4. 完善加载等待逻辑

在打开详情页后,增加页面加载等待,比如判断标题是否加载完成:

wait.until(EC.title_contains("Exhibitor"))

5. 补充字段提取与异常处理

为避免数据错位,异常时填充空值,同时补充官网、LinkedIn等字段的稳定提取逻辑:

# 提取官网链接示例
try:
    website = driver.find_element(By.XPATH, "//a[contains(@href, 'http') and not(contains(@href, 'linkedin'))]").get_attribute('href')
except:
    website = None
exhibitor_web.append(website)

完整修复后代码示例

exhibitor_names = []
exhibitor_info = []
exhibitor_web = []
exhibitor_linkedin = []
exhibitor_twitter = []
exhibitor_email = []
exhibitor_contact = []
exhibitor_booth = []

import time
import logging
import pandas as pd
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# Configure logging
logging.basicConfig(level=logging.INFO)
logger = logging.getLogger()

# Initialize WebDriver and Options
chrome_options = Options()
chrome_options.add_argument("--start-maximized")
driver = webdriver.Chrome(options=chrome_options)
wait = WebDriverWait(driver, 15)

# 主页面加载与iframe切换
driver.get("https://asiatechxsg.com/exhibitors/")
wait.until(EC.frame_to_be_available_and_switch_to_it((By.CLASS_NAME, "IframeModule_iframe__JCvXg")))

# 滚动加载所有参展商
while True:
    prev_height = driver.execute_script("return document.body.scrollHeight")
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(2)
    new_height = driver.execute_script("return document.body.scrollHeight")
    last_element = driver.find_elements(By.XPATH, "//img[@alt='Singapore Centre for Social Enterprise, raiSE Ltd']")
    if len(last_element) >= 1 or new_height == prev_height:
        break

# 提取所有参展商链接URL
exhibitor_links = [link.get_attribute('href') for link in driver.find_elements(By.XPATH, "//a[contains(@href, '/widget/event/asia-tech-x-singapore-2024/exhibitor/')]")]
logger.info(f"Found {len(exhibitor_links)} exhibitor links.")

# 遍历每个参展商页面
for index, url in enumerate(exhibitor_links):
    try:
        driver.get(url)
        logger.info(f"Processing exhibitor {index + 1}/{len(exhibitor_links)}")
        
        # 切换到详情页的iframe
        wait.until(EC.frame_to_be_available_and_switch_to_it((By.CLASS_NAME, "IframeModule_iframe__JCvXg")))
        
        # 提取参展商名称
        name = wait.until(EC.visibility_of_element_located((By.TAG_NAME, "h1"))).text
        exhibitor_names.append(name)
        
        # 提取详情描述
        info_div = wait.until(EC.visibility_of_element_located((By.XPATH, "//div[contains(@class, 'sc-dd6f9f7c-0')]")))
        paragraphs = info_div.find_elements(By.TAG_NAME, "p")
        info_text = '\n'.join([p.text.strip() for p in paragraphs if p.text.strip()])
        exhibitor_info.append(info_text)
        
        # 提取官网链接
        try:
            web = driver.find_element(By.XPATH, "//a[contains(@href, 'http') and not(contains(@href, 'linkedin')) and not(contains(@href, 'twitter'))]").get_attribute('href')
        except:
            web = None
        exhibitor_web.append(web)
        
        # 提取LinkedIn链接
        try:
            linkedin = driver.find_element(By.XPATH, "//a[contains(@href, 'linkedin')]").get_attribute('href')
        except:
            linkedin = None
        exhibitor_linkedin.append(linkedin)
        
        # 提取展位号
        try:
            booth = driver.find_element(By.XPATH, "//div[contains(text(), 'Booth')]/following-sibling::div").text
        except:
            booth = None
        exhibitor_booth.append(booth)
        
        logger.info(f"Successfully extracted data for {name}")
        
    except Exception as e:
        logger.error(f"Failed to process exhibitor {index + 1}: {str(e)}")
        # 填充空值避免数据错位
        exhibitor_names.append(None)
        exhibitor_info.append(None)
        exhibitor_web.append(None)
        exhibitor_linkedin.append(None)
        exhibitor_booth.append(None)
        continue

# 关闭浏览器
driver.quit()

# 创建DataFrame并导出CSV
exhibitors_df = pd.DataFrame({
    'Exhibitor Name': exhibitor_names,
    'Description': exhibitor_info,
    'Website': exhibitor_web,
    'LinkedIn': exhibitor_linkedin,
    'Booth Number': exhibitor_booth
})

# 过滤空值记录
exhibitors_df = exhibitors_df.dropna(subset=['Exhibitor Name'])

# 导出CSV
exhibitors_df.to_csv('asiatechx_exhibitors.csv', index=False, encoding='utf-8-sig')
logger.info(f"Data exported to asiatechx_exhibitors.csv, total records: {len(exhibitors_df)}")

注意事项

  • 动态类名可能随时变化,每次运行前需用浏览器开发者工具确认元素选择器
  • 可适当调整等待超时时间或增加time.sleep(),适配网络加载速度
  • 建议添加请求间隔(如每次页面跳转后等待1-2秒),避免触发网站反爬限制

内容的提问来源于stack exchange,提问作者user22279494

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 02:45:04