You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium多下拉框组合选择报错及批量遍历方案咨询

问题:遍历多下拉框选项组合的网络爬虫实现

我正在构建一个网络爬虫,需要遍历5个下拉框的所有选项组合,针对每个组合点击按钮加载数据页面,并将数据存储到字典中。

目标网站:http://siops.datasus.gov.br/filtro_rel_ges_covid_municipal.php?S=1&UF=12;&Municipio=120001;&Ano=2020&Periodo=20

当前测试代码如下:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import Select

# General Stuff about the website
path = '/Users/admin/desktop/projects/scraper/chromedriver'
options = Options()
options.headless = True
options.add_argument("--window-size=1920,1200")
driver = webdriver.Chrome(options=options, executable_path=path)
website = 'http://siops.datasus.gov.br/filtro_rel_ges_covid_municipal.php'
driver.get(website)

# Initial Test: printing the title
print(driver.title)
print()

# Dictionary to Store stuff in
totals = {}

# Drop Down Menus
year_select = Select(driver.find_element(By.XPATH, '//*[@id="cmbAno"]'))
uf_select = Select(driver.find_element(By.XPATH, '//*[@id="cmbUF"]'))

### THIS IS WHERE THE ERROR IS OCCURING ###
# Choose from the drop down menus 
uf_select.select_by_value('29')
year_select.select_by_value('2020')

# Submit button on the page
submit_button = driver.find_element(By.XPATH, '//*[@id="container"]/div[2]/form/div[2]/div/input[2]')
submit_button.click()

# Pulling data from the webpage
nameof = driver.find_element(By.XPATH, '//*[@id="arearelatorio"]/div[1]/div/table[1]/tbody/tr[2]').text
total_balance = driver.find_element(By.XPATH, '//*[@id="arearelatorio"]/div[1]/div/table[3]/tbody/tr[9]/td[2]').text
paid_expenses = driver.find_element(By.XPATH, '//*[@id="arearelatorio"]/div[1]/div/table[4]/tbody/tr[11]/td[4]').text

# Update Dictionary with the new info
totals.update({nameof: [total_balance, paid_expenses]})
totals.update({'this is a test': ['testing stuff']})

# Print the final Dictionary and quit
print(totals)
driver.quit()

遇到的问题:当同时选择UF下拉框值29和年份下拉框值2020这一组合时,代码运行失败,但单独选择其中一个下拉框时代码可正常执行。现咨询如何实现单次迭代中尝试多个下拉框选项的所有组合?


解决方案

一、解决双下拉框同时选择失败的问题

1. 添加显式等待

网站在选择第一个下拉框后,可能需要动态加载第二个下拉框的可选状态,直接选择会导致元素交互异常。引入WebDriverWait等待元素可交互:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 选择UF后,等待年份下拉框可点击
uf_select.select_by_value('29')
WebDriverWait(driver, 10).until(
    EC.element_to_be_clickable((By.XPATH, '//*[@id="cmbAno"]'))
)
year_select.select_by_value('2020')

2. 避免失效的下拉框实例

部分网站会在选择下拉框后刷新DOM,之前获取的Select实例可能失效,建议每次选择前重新定位元素:

# 选择UF
Select(driver.find_element(By.XPATH, '//*[@id="cmbUF"]')).select_by_value('29')
# 等待年份下拉框重新加载完成
WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.XPATH, '//*[@id="cmbAno"]'))
)
# 重新获取年份下拉框实例并选择
Select(driver.find_element(By.XPATH, '//*[@id="cmbAno"]')).select_by_value('2020')

二、遍历所有下拉框选项组合的实现

针对网站的联动下拉框(如选择UF后才加载Municipio选项),采用动态嵌套循环的方式遍历所有组合:

1. 定义工具函数获取有效选项

先写一个函数,提取下拉框中带实际值的选项(跳过无意义的提示选项):

def get_valid_options(select_xpath):
    select_element = Select(driver.find_element(By.XPATH, select_xpath))
    # 返回(value, text)格式的有效选项列表
    return [(opt.get_attribute('value'), opt.text) 
            for opt in select_element.options 
            if opt.get_attribute('value') and opt.text.strip()]

2. 嵌套循环遍历所有组合

使用itertools.product简化多下拉框的组合遍历,同时处理联动下拉框的动态加载:

from itertools import product
import time

# 初始化存储字典
totals = {}

# 获取非联动下拉框的初始选项(如年份、周期)
ano_options = get_valid_options('//*[@id="cmbAno"]')
periodo_options = get_valid_options('//*[@id="cmbPeriodo"]')

# 遍历UF选项(触发Municipio联动加载)
for uf_val, uf_text in get_valid_options('//*[@id="cmbUF"]'):
    # 选择当前UF
    Select(driver.find_element(By.XPATH, '//*[@id="cmbUF"]')).select_by_value(uf_val)
    # 等待Municipio下拉框加载完成
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.XPATH, '//*[@id="cmbMunicipio"]'))
    )
    # 获取当前UF对应的Municipio选项
    municipio_options = get_valid_options('//*[@id="cmbMunicipio"]')
    
    # 遍历年份、周期的组合
    for (ano_val, ano_text), (periodo_val, periodo_text) in product(ano_options, periodo_options):
        # 选择年份
        Select(driver.find_element(By.XPATH, '//*[@id="cmbAno"]')).select_by_value(ano_val)
        # 选择周期
        Select(driver.find_element(By.XPATH, '//*[@id="cmbPeriodo"]')).select_by_value(periodo_val)
        
        # 遍历当前UF下的所有Municipio选项
        for mun_val, mun_text in municipio_options:
            # 选择Municipio
            Select(driver.find_element(By.XPATH, '//*[@id="cmbMunicipio"]')).select_by_value(mun_val)
            
            # 等待提交按钮可点击并点击
            submit_btn = WebDriverWait(driver, 10).until(
                EC.element_to_be_clickable((By.XPATH, '//*[@id="container"]/div[2]/form/div[2]/div/input[2]'))
            )
            submit_btn.click()
            
            # 等待数据页面加载完成
            WebDriverWait(driver, 15).until(
                EC.presence_of_element_located((By.XPATH, '//*[@id="arearelatorio"]'))
            )
            
            # 提取数据并存储(添加异常处理避免崩溃)
            try:
                entity_name = driver.find_element(By.XPATH, '//*[@id="arearelatorio"]/div[1]/div/table[1]/tbody/tr[2]').text
                total_balance = driver.find_element(By.XPATH, '//*[@id="arearelatorio"]/div[1]/div/table[3]/tbody/tr[9]/td[2]').text
                paid_expenses = driver.find_element(By.XPATH, '//*[@id="arearelatorio"]/div[1]/div/table[4]/tbody/tr[11]/td[4]').text
                
                # 用组合信息作为字典键,方便后续查询
                key = f"{uf_text}-{mun_text}-{ano_text}-{periodo_text}"
                totals[key] = [total_balance, paid_expenses]
                print(f"已存储组合: {key}")
            except Exception as e:
                print(f"组合 [{uf_val}, {mun_val}, {ano_val}, {periodo_val}] 数据提取失败: {str(e)}")
            
            # 返回筛选页面,准备下一组组合
            driver.back()
            # 等待筛选页面元素加载完成
            WebDriverWait(driver, 10).until(
                EC.presence_of_element_located((By.XPATH, '//*[@id="cmbUF"]'))
            )
            # 短暂停顿避免请求过快
            time.sleep(0.5)

# 打印最终结果
print(totals)
driver.quit()

3. 关键优化点

  • 禁用Headless调试:如果仍出现异常,先关闭options.headless = True,直观观察页面交互过程,定位问题根源。
  • 全局异常捕获:对每个步骤的元素操作添加try-except,避免单个组合失败导致整个爬虫终止。
  • 请求频率控制:添加time.sleep()或使用随机延迟,降低被网站反爬机制拦截的风险。

内容的提问来源于stack exchange,提问作者Seaborg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 21:45:38