Selenium爬取参展商页面仅获单个标题 如何抓取全部标题
问题分析
当前代码只能抓取1个标题的核心原因有两个:
- 定位
div.card-exhibitor时默认只返回匹配到的第一个元素,点击后仅进入单个参展商详情页抓取h1标题 - 目标参展商列表是滚动懒加载模式,初始页面加载完成后只渲染首批卡片,剩余内容需要滚动到页面底部才会逐批加载,现有代码没有做全量加载的处理
全量抓取修改思路
- 放弃逐个点击进入详情页抓标题的逻辑:参展商名称直接在列表卡片上就有展示,直接在列表页提取即可,大幅提升爬取效率
- 增加滚动懒加载处理:循环滚动到页面底部,等待新内容加载,直到页面高度不再变化,确认所有参展商卡片渲染完成
- 用
find_elements方法获取所有匹配的参展商卡片元素,遍历提取每个卡片内的标题文本
修改后完整可运行代码
import time from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from webdriver_manager.chrome import ChromeDriverManager options = webdriver.ChromeOptions() options.add_argument("--headless") options.add_argument("--no-sandbox") options.add_argument("--disable-gpu") options.add_argument("--window-size=1920x1080") options.add_argument("--disable-extensions") chrome_driver = webdriver.Chrome( service=Service(ChromeDriverManager().install()), options=options ) def supplyvan_scraper(): with chrome_driver as driver: driver.implicitly_wait(15) URL = 'https://www.eurobike.com/en/index-exhibitors/exhibitors/?' driver.get(URL) time.sleep(3) # 循环滚动加载所有参展商 last_height = driver.execute_script("return document.body.scrollHeight") while True: # 滚动到页面底部 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # 等待新内容加载 time.sleep(2) # 对比页面高度判断是否加载完成 new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height # 遍历所有卡片提取标题 exhibitor_cards = driver.find_elements(By.CSS_SELECTOR, "div.card-exhibitor") all_titles = [] for card in exhibitor_cards: title = card.find_element(By.CSS_SELECTOR, "h3.card-title").text.strip() if title: all_titles.append(title) # 输出结果 print(f"共抓取到{len(all_titles)}个参展商标题:") for t in all_titles: print(t) supplyvan_scraper()
补充说明
- 如果需要进入每个参展商详情页抓取更完整的信息,可以在拿到所有卡片的跳转链接后逐个访问,注意控制访问频率,避免被站点反爬拦截
- 滚动等待的
time.sleep(2)可以根据实际网络情况调整,网络较差时可以适当调大等待时长,避免新内容未加载完成就判定加载结束
内容的提问来源于stack exchange,提问作者Amen Aziz
相关产品推荐
相关产品推荐

