使用Selenium的Python网页爬虫循环迭代若干次后停止工作
问题描述
- 使用Python Selenium爬取https://www.pic.int网站的表格数据,通过循环遍历下拉菜单中的所有国家获取数据
- 前约10次迭代(通常到巴林)运行正常,之后无法获取国家名称,输出为空字符串
- 怀疑和第9个国家出现的弹窗有关,点击
confirm_no.k-button关闭弹窗后,下一次循环出现问题 - 曾尝试直接点击下拉选项无效,改用
send_keys方法输入国家名称
故障原因分析
- 弹窗处理逻辑错误:当前代码在关闭弹窗后使用
continue直接进入下一次循环,导致当前国家的表格数据未被爬取,同时下拉菜单的元素状态可能因弹窗干扰失效 - 过时元素问题:弹窗关闭后页面DOM可能重新渲染,之前获取的
dropdown_options列表中的元素已变为stale element(过时元素),无法再获取文本 - BeautifulSoup语法错误:
rows = relevant_section.find_all("tr", 'class'==["row1", "row2"])的写法错误,这不是BeautifulSoup查找类的正确方式
解决方案
针对上述问题,调整后的代码如下:
from selenium import webdriver from selenium.webdriver.chrome.service import Service as ChromeService from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.keys import Keys import itertools from bs4 import BeautifulSoup # 初始化存储列表 ChemName = [] Category = [] Country = [] Response = [] Decision = [] Date = [] driver = webdriver.Chrome(service=ChromeService(ChromeDriverManager().install())) driver.get("https://www.pic.int/Procedures/ImportResponses/Database/tabid/1370/language/en-US/Default.aspx") driver.implicitly_wait(10) wait = WebDriverWait(driver, 10) # 切换到"Import Responses by Party"标签页 party_tab = wait.until(EC.element_to_be_clickable((By.XPATH, "//*[@aria-controls = 'tabstrip_ICR-2']"))) party_tab.click() # 等待标签页内容加载完成 wait.until(EC.presence_of_element_located((By.CLASS_NAME, "byParty.k-content.k-state-active"))) # 预提取所有国家名称,避免循环中出现过时元素问题 country_dropdown = wait.until(EC.element_to_be_clickable((By.CLASS_NAME, "k-icon.k-i-arrow-s"))) country_dropdown.click() dropdown_options = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "ul#ddlParty_listbox li.k-item"))) # 过滤空文本,提取有效国家名称 country_list = [opt.text.strip() for opt in dropdown_options if opt.text.strip()] # 关闭下拉框 country_dropdown.click() # 遍历国家列表 for country_name in itertools.islice(country_list, 0, 12): print(f"当前处理国家: {country_name}") # 定位下拉输入框,输入国家名称 input_box = wait.until(EC.element_to_be_clickable((By.XPATH, '//*[@aria-owns="ddlParty_listbox"]'))) input_box.click() input_box.send_keys(Keys.CONTROL + 'a') input_box.send_keys(Keys.DELETE) input_box.send_keys(country_name) input_box.send_keys(Keys.ENTER) # 处理EU弹窗 try: popup = wait.until(EC.presence_of_element_located((By.ID, "windowEU"))) confirm_btn = wait.until(EC.element_to_be_clickable((By.CLASS_NAME, "confirm_no.k-button"))) confirm_btn.click() # 等待弹窗完全关闭,确保页面稳定 wait.until(EC.staleness_of(popup)) except: pass # 等待目标表格加载完成 wait.until(EC.presence_of_element_located((By.ID, "IRview"))) # 解析表格数据 page_source = driver.page_source soup = BeautifulSoup(page_source, "html.parser") relevant_section = soup.find("table", id="IRview") # 修正BeautifulSoup的类查找语法 rows = relevant_section.find_all("tr", class_=["row1", "row2"]) if rows: for row in rows: columns = row.find_all("td") # 确保存在6列数据再提取 if len(columns) >= 6: chemical_name = columns[0].text.strip() category = columns[1].text.strip() party = columns[2].text.strip() resp = columns[3].text.strip() dec = columns[4].text.strip() dat = columns[5].text.strip() ChemName.append(chemical_name) Category.append(category) Country.append(party) Response.append(resp) Decision.append(dec) Date.append(dat) # 关闭浏览器 driver.quit()
关键调整说明
- 预提取国家列表:提前将所有国家名称提取到列表中,避免循环中使用已失效的DOM元素
- 修复弹窗处理逻辑:去掉
continue,弹窗关闭后继续完成当前国家的数据爬取,同时等待弹窗元素过时确保页面稳定 - 修正BeautifulSoup语法:将错误的类查找方式改为
class_=["row1", "row2"] - 增强等待逻辑:使用
staleness_of确保弹窗完全关闭后再进行后续操作 - 调整列数判断:从
>=4改为>=6,保证能获取到全部6列数据
内容的提问来源于stack exchange,提问作者lcasse
相关产品推荐
相关产品推荐

