You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium抓取MGA网站Licensee Name下拉列表并逐个搜索?

问题分析与解决

你的代码存在几个核心问题:

  • 用了动态生成的aria-activedescendant值定位元素,这个ID每次页面加载都会变化,无法稳定定位目标
  • 错误调用了不存在的.input()方法,WebElement对象没有这个属性
  • 导入大量无关库(比如js、numpy、ctypes),徒增代码复杂度
  • 无头模式参数格式错误,新版Chrome需用--headless=new
  • 隐式等待设置过长,且未结合显式等待处理动态加载元素

下面是针对你需求的完整可行代码,核心逻辑是处理Select2动态下拉组件,获取所有选项后逐个执行搜索:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.chrome.options import Options
import time
import pandas as pd

# 初始化浏览器配置
opts = Options()
prefs = {"profile.default_content_setting_values.notifications": 2}
opts.add_experimental_option("prefs", prefs)
# 新版Chrome无头模式参数
opts.add_argument('--headless=new')
# 禁用图片加载提升爬取速度
opts.add_argument('--blink-settings=imagesEnabled=false')

# 注意:ChromeDriver路径需与你的Chrome版本匹配,推荐使用webdriver-manager自动管理驱动
driver = webdriver.Chrome('C:/chromedriver_win32/chromedriver.exe', options=opts)
driver.maximize_window()

url = 'https://www.mga.org.mt/licences/'
driver.get(url)

# 显式等待页面加载,定位Licensee Name的Select2搜索输入框
wait = WebDriverWait(driver, 20)
licensee_input = wait.until(EC.presence_of_element_located(
    (By.CSS_SELECTOR, 'span.select2-container[id*="LicenseeName"] input.select2-search__field')
))

# 点击输入框展开下拉列表
licensee_input.click()
time.sleep(1)  # 等待下拉选项加载

# 获取所有下拉选项元素
options = wait.until(EC.presence_of_all_elements_located(
    (By.CSS_SELECTOR, 'ul.select2-results__options li.select2-results__option')
))

# 提取有效选项文本(跳过空的占位选项)
option_texts = [opt.text.strip() for opt in options if opt.text.strip()]

# 存储搜索结果的列表
results = []

for text in option_texts:
    try:
        # 重新激活输入框(避免下拉收起)
        licensee_input.click()
        time.sleep(0.5)
        # 清空输入框并输入当前选项文本
        licensee_input.clear()
        licensee_input.send_keys(text)
        time.sleep(1)
        # 选择过滤后的匹配选项
        matched_option = wait.until(EC.element_to_be_clickable(
            (By.CSS_SELECTOR, 'ul.select2-results__options li.select2-results__option--highlighted')
        ))
        matched_option.click()
        time.sleep(0.5)
        
        # 点击搜索按钮
        search_btn = wait.until(EC.element_to_be_clickable(
            (By.ID, 'btnSearch')
        ))
        search_btn.click()
        time.sleep(2)  # 等待搜索结果加载
        
        # 提取搜索结果表格数据(可根据需求修改提取逻辑)
        soup = BeautifulSoup(driver.page_source, 'html.parser')
        table = soup.find('table', id='tblLicences')
        if table:
            rows = table.find_all('tr')[1:]  # 跳过表头行
            for row in rows:
                cols = [col.text.strip() for col in row.find_all('td')]
                results.append({
                    'Licensee Name': text,
                    'Licence Number': cols[0],
                    'Licence Type': cols[1],
                    'Status': cols[2],
                    'Issue Date': cols[3],
                    'Expiry Date': cols[4]
                })
        
        # 回到搜索页面,准备下一次搜索
        driver.get(url)
        licensee_input = wait.until(EC.presence_of_element_located(
            (By.CSS_SELECTOR, 'span.select2-container[id*="LicenseeName"] input.select2-search__field')
        ))
        
    except Exception as e:
        print(f"处理选项 {text} 时出错: {str(e)}")
        continue

# 将结果保存为CSV文件
df = pd.DataFrame(results)
df.to_csv('mga_licences.csv', index=False, encoding='utf-8-sig')

# 关闭浏览器
driver.quit()

关键注意事项:

  1. Select2组件处理:这类动态下拉不能用普通<select>元素定位,必须通过其专属的搜索框、下拉选项CSS选择器操作
  2. 显式等待优先:用WebDriverWait替代过长的隐式等待,确保元素加载完成后再执行操作
  3. 反爬规避:添加适当的time.sleep()避免请求过快被拦截,也可考虑添加随机延迟
  4. 驱动版本匹配:确保ChromeDriver版本与Chrome浏览器版本完全一致,否则会启动失败,推荐用webdriver-manager库自动管理驱动

内容的提问来源于stack exchange,提问作者Mohamed Hedeya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 15:40:48