You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium的Python网页爬虫循环迭代若干次后停止工作

问题描述
  • 使用Python Selenium爬取https://www.pic.int网站的表格数据,通过循环遍历下拉菜单中的所有国家获取数据
  • 前约10次迭代(通常到巴林)运行正常,之后无法获取国家名称,输出为空字符串
  • 怀疑和第9个国家出现的弹窗有关,点击confirm_no.k-button关闭弹窗后,下一次循环出现问题
  • 曾尝试直接点击下拉选项无效,改用send_keys方法输入国家名称
故障原因分析
  1. 弹窗处理逻辑错误:当前代码在关闭弹窗后使用continue直接进入下一次循环,导致当前国家的表格数据未被爬取,同时下拉菜单的元素状态可能因弹窗干扰失效
  2. 过时元素问题:弹窗关闭后页面DOM可能重新渲染,之前获取的dropdown_options列表中的元素已变为stale element(过时元素),无法再获取文本
  3. BeautifulSoup语法错误:rows = relevant_section.find_all("tr", 'class'==["row1", "row2"])的写法错误,这不是BeautifulSoup查找类的正确方式
解决方案

针对上述问题,调整后的代码如下:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service as ChromeService
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.keys import Keys
import itertools
from bs4 import BeautifulSoup

# 初始化存储列表
ChemName = []
Category = []
Country = []
Response = []
Decision = []
Date = []

driver = webdriver.Chrome(service=ChromeService(ChromeDriverManager().install()))
driver.get("https://www.pic.int/Procedures/ImportResponses/Database/tabid/1370/language/en-US/Default.aspx")    
driver.implicitly_wait(10)
wait = WebDriverWait(driver, 10)

# 切换到"Import Responses by Party"标签页
party_tab = wait.until(EC.element_to_be_clickable((By.XPATH, "//*[@aria-controls = 'tabstrip_ICR-2']")))
party_tab.click()

# 等待标签页内容加载完成
wait.until(EC.presence_of_element_located((By.CLASS_NAME, "byParty.k-content.k-state-active")))

# 预提取所有国家名称,避免循环中出现过时元素问题
country_dropdown = wait.until(EC.element_to_be_clickable((By.CLASS_NAME, "k-icon.k-i-arrow-s")))
country_dropdown.click()
dropdown_options = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "ul#ddlParty_listbox li.k-item")))
# 过滤空文本,提取有效国家名称
country_list = [opt.text.strip() for opt in dropdown_options if opt.text.strip()]
# 关闭下拉框
country_dropdown.click()

# 遍历国家列表
for country_name in itertools.islice(country_list, 0, 12):
    print(f"当前处理国家: {country_name}")
    
    # 定位下拉输入框,输入国家名称
    input_box = wait.until(EC.element_to_be_clickable((By.XPATH, '//*[@aria-owns="ddlParty_listbox"]')))
    input_box.click()
    input_box.send_keys(Keys.CONTROL + 'a')
    input_box.send_keys(Keys.DELETE)
    input_box.send_keys(country_name)
    input_box.send_keys(Keys.ENTER)
    
    # 处理EU弹窗
    try:
        popup = wait.until(EC.presence_of_element_located((By.ID, "windowEU")))
        confirm_btn = wait.until(EC.element_to_be_clickable((By.CLASS_NAME, "confirm_no.k-button")))
        confirm_btn.click()
        # 等待弹窗完全关闭,确保页面稳定
        wait.until(EC.staleness_of(popup))
    except:
        pass
    
    # 等待目标表格加载完成
    wait.until(EC.presence_of_element_located((By.ID, "IRview")))
    
    # 解析表格数据
    page_source = driver.page_source
    soup = BeautifulSoup(page_source, "html.parser")
    relevant_section = soup.find("table", id="IRview")
    
    # 修正BeautifulSoup的类查找语法
    rows = relevant_section.find_all("tr", class_=["row1", "row2"])
    
    if rows:
        for row in rows:
            columns = row.find_all("td")
            # 确保存在6列数据再提取
            if len(columns) >= 6:
                chemical_name = columns[0].text.strip()
                category = columns[1].text.strip()
                party = columns[2].text.strip()
                resp = columns[3].text.strip()
                dec = columns[4].text.strip()
                dat = columns[5].text.strip()
                
                ChemName.append(chemical_name)
                Category.append(category)
                Country.append(party)
                Response.append(resp)
                Decision.append(dec)
                Date.append(dat)

# 关闭浏览器
driver.quit()

关键调整说明

  • 预提取国家列表:提前将所有国家名称提取到列表中,避免循环中使用已失效的DOM元素
  • 修复弹窗处理逻辑:去掉continue,弹窗关闭后继续完成当前国家的数据爬取,同时等待弹窗元素过时确保页面稳定
  • 修正BeautifulSoup语法:将错误的类查找方式改为class_=["row1", "row2"]
  • 增强等待逻辑:使用staleness_of确保弹窗完全关闭后再进行后续操作
  • 调整列数判断:从>=4改为>=6,保证能获取到全部6列数据

内容的提问来源于stack exchange,提问作者lcasse

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 00:06:24