You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium批量打开网页下拉菜单并获取子分类href

解决方案

核心思路

网站下拉菜单为动态触发模式(鼠标悬停后展开),需先模拟用户行为触发菜单显示,再定位子分类元素提取href。具体步骤:

  • 定位所有顶层分类的触发元素
  • 对每个触发元素执行鼠标悬停操作,展开下拉菜单
  • 等待子菜单加载完成后,提取所有子分类的链接
  • 统一收集并去重所有子分类href

修改后的代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.action_chains import ActionChains

URL = "https://albiononline2d.com/en/item"

driver = webdriver.Chrome()
wait = WebDriverWait(driver, 10)
driver.get(URL)

def get_all_subcategory_hrefs():
    """获取所有子分类的href链接"""
    subcategory_hrefs = []
    
    # 定位所有带下拉功能的顶层分类触发元素
    top_categories = wait.until(
        EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'a.dropdown-toggle'))
    )
    
    for category in top_categories:
        # 鼠标悬停展开下拉菜单
        ActionChains(driver).move_to_element(category).perform()
        
        # 通过顶层元素的aria-controls属性,精准定位对应子菜单
        sub_menu_id = category.get_attribute('aria-controls')
        sub_menu = wait.until(
            EC.visibility_of_element_located((By.ID, sub_menu_id))
        )
        sub_links = sub_menu.find_elements(By.TAG_NAME, 'a')
        
        # 提取子分类href并去重
        for link in sub_links:
            href = link.get_attribute('href')
            if href and href not in subcategory_hrefs:
                subcategory_hrefs.append(href)
    
    return subcategory_hrefs

# 调用函数获取所有子分类链接
all_subcategory_links = get_all_subcategory_hrefs()
for link in all_subcategory_links:
    print(link)

driver.quit()

代码说明

  • 用WebDriverWait显式等待元素加载,适配动态页面的延迟加载特性,比隐式等待更可靠
  • 通过aria-controls属性关联顶层分类与子菜单,避免因类名变化导致的定位失效
  • 用ActionChains模拟鼠标悬停,还原用户触发下拉菜单的真实操作
  • 加入去重逻辑,避免重复收集相同的子分类链接

子分类页面物品抓取示例

拿到子分类链接后,可遍历进入每个页面抓取物品信息,示例逻辑如下:

def scrape_subcategory_items(subcategory_href):
    driver.get(subcategory_href)
    # 等待物品卡片加载完成
    item_cards = wait.until(
        EC.presence_of_all_elements_located((By.CLASS_NAME, 'item-card'))
    )
    for card in item_cards:
        # 提取物品名称、图标、属性等信息
        item_name = card.find_element(By.CLASS_NAME, 'item-name').text
        item_icon = card.find_element(By.TAG_NAME, 'img').get_attribute('src')
        print(f"物品名称:{item_name},图标链接:{item_icon}")

# 遍历所有子分类页面抓取信息
for href in all_subcategory_links:
    scrape_subcategory_items(href)

内容的提问来源于stack exchange,提问作者Gabrielgad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 07:47:16