如何用Selenium抓取MGA网站Licensee Name下拉列表并逐个搜索?
问题分析与解决
你的代码存在几个核心问题:
- 用了动态生成的
aria-activedescendant值定位元素,这个ID每次页面加载都会变化,无法稳定定位目标 - 错误调用了不存在的
.input()方法,WebElement对象没有这个属性 - 导入大量无关库(比如
js、numpy、ctypes),徒增代码复杂度 - 无头模式参数格式错误,新版Chrome需用
--headless=new - 隐式等待设置过长,且未结合显式等待处理动态加载元素
下面是针对你需求的完整可行代码,核心逻辑是处理Select2动态下拉组件,获取所有选项后逐个执行搜索:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.chrome.options import Options import time import pandas as pd # 初始化浏览器配置 opts = Options() prefs = {"profile.default_content_setting_values.notifications": 2} opts.add_experimental_option("prefs", prefs) # 新版Chrome无头模式参数 opts.add_argument('--headless=new') # 禁用图片加载提升爬取速度 opts.add_argument('--blink-settings=imagesEnabled=false') # 注意:ChromeDriver路径需与你的Chrome版本匹配,推荐使用webdriver-manager自动管理驱动 driver = webdriver.Chrome('C:/chromedriver_win32/chromedriver.exe', options=opts) driver.maximize_window() url = 'https://www.mga.org.mt/licences/' driver.get(url) # 显式等待页面加载,定位Licensee Name的Select2搜索输入框 wait = WebDriverWait(driver, 20) licensee_input = wait.until(EC.presence_of_element_located( (By.CSS_SELECTOR, 'span.select2-container[id*="LicenseeName"] input.select2-search__field') )) # 点击输入框展开下拉列表 licensee_input.click() time.sleep(1) # 等待下拉选项加载 # 获取所有下拉选项元素 options = wait.until(EC.presence_of_all_elements_located( (By.CSS_SELECTOR, 'ul.select2-results__options li.select2-results__option') )) # 提取有效选项文本(跳过空的占位选项) option_texts = [opt.text.strip() for opt in options if opt.text.strip()] # 存储搜索结果的列表 results = [] for text in option_texts: try: # 重新激活输入框(避免下拉收起) licensee_input.click() time.sleep(0.5) # 清空输入框并输入当前选项文本 licensee_input.clear() licensee_input.send_keys(text) time.sleep(1) # 选择过滤后的匹配选项 matched_option = wait.until(EC.element_to_be_clickable( (By.CSS_SELECTOR, 'ul.select2-results__options li.select2-results__option--highlighted') )) matched_option.click() time.sleep(0.5) # 点击搜索按钮 search_btn = wait.until(EC.element_to_be_clickable( (By.ID, 'btnSearch') )) search_btn.click() time.sleep(2) # 等待搜索结果加载 # 提取搜索结果表格数据(可根据需求修改提取逻辑) soup = BeautifulSoup(driver.page_source, 'html.parser') table = soup.find('table', id='tblLicences') if table: rows = table.find_all('tr')[1:] # 跳过表头行 for row in rows: cols = [col.text.strip() for col in row.find_all('td')] results.append({ 'Licensee Name': text, 'Licence Number': cols[0], 'Licence Type': cols[1], 'Status': cols[2], 'Issue Date': cols[3], 'Expiry Date': cols[4] }) # 回到搜索页面,准备下一次搜索 driver.get(url) licensee_input = wait.until(EC.presence_of_element_located( (By.CSS_SELECTOR, 'span.select2-container[id*="LicenseeName"] input.select2-search__field') )) except Exception as e: print(f"处理选项 {text} 时出错: {str(e)}") continue # 将结果保存为CSV文件 df = pd.DataFrame(results) df.to_csv('mga_licences.csv', index=False, encoding='utf-8-sig') # 关闭浏览器 driver.quit()
关键注意事项:
- Select2组件处理:这类动态下拉不能用普通
<select>元素定位,必须通过其专属的搜索框、下拉选项CSS选择器操作 - 显式等待优先:用
WebDriverWait替代过长的隐式等待,确保元素加载完成后再执行操作 - 反爬规避:添加适当的
time.sleep()避免请求过快被拦截,也可考虑添加随机延迟 - 驱动版本匹配:确保ChromeDriver版本与Chrome浏览器版本完全一致,否则会启动失败,推荐用
webdriver-manager库自动管理驱动
内容的提问来源于stack exchange,提问作者Mohamed Hedeya
相关产品推荐
相关产品推荐

