You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效抓取单网站海量URL?Selenium优化及重试方案咨询

优化10000+页面的Selenium URL提取方案

现有Selenium代码可遍历网站分页提取URL,但面对10000+页面时耗时极长;需要实现页面点击失败3次后跳过该页继续下一页的逻辑,且网站无公开API。

关键优化措施

1. 用显式等待替代固定休眠

固定time.sleep()会浪费大量等待时间,改用WebDriverWait等待目标元素加载完成,仅在必要时等待,大幅提升遍历效率。

2. 启用无头浏览器模式

禁用Chrome的UI渲染,减少资源占用,运行速度至少提升30%以上。

3. 优化文件IO操作

原代码每次循环都重写整个文件,改为追加写入或定时批量写入,避免重复IO开销。

4. 修复分页重试逻辑

原代码重试失败后会终止循环,调整为:3次重试失败后,直接尝试跳转到下一页(优先通过URL构造,避免依赖分页按钮),确保遍历不中断。

5. 精简元素定位逻辑

复用定位表达式,避免重复编写,同时使用更精准的定位方式(比如相对XPath)减少元素查找失败概率。


改进后的完整代码

import random
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from selenium.common.exceptions import WebDriverException, NoSuchElementException, TimeoutException

user_agents = [
    "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/109.0.0.0 Safari/537.36",
    "Mozilla/5.0 (Macintosh; Intel Mac OS X 13_1) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/16.1 Safari/605.1.15"
]

# 配置Chrome选项
chrome_options = Options()
chrome_options.add_argument(f'user-agent={random.choice(user_agents)}')
chrome_options.add_argument('--headless=new')  # 无头模式
chrome_options.add_argument('--disable-gpu')
chrome_options.add_argument('--no-sandbox')
chrome_options.add_argument('--disable-dev-shm-usage')

driver = webdriver.Chrome(options=chrome_options)
wait = WebDriverWait(driver, 10)  # 显式等待超时时间
main_url = 'https://sekolah.data.kemdikbud.go.id/index.php/Chome/pencarian/'
href_set = set()
page = 1
max_retries = 3
# 提前定义定位表达式
PAGELOAD_SECTION = (By.ID, 'pageload')
VIEW_LINKS = (By.XPATH, './/li[@class="list-group-item"]/a[contains(., "Lihat")]')  # 相对路径定位
NEXT_PAGE_TEMPLATE = 'https://sekolah.data.kemdikbud.go.id/index.php/Chome/pencarian/{}'

try:
    driver.get(main_url)
    # 等待搜索按钮并点击
    search_button = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, 'button[type="submit"]')))
    search_button.click()

    while True:
        success = False
        for attempt in range(max_retries):
            try:
                # 等待页面内容加载
                section = wait.until(EC.presence_of_element_located(PAGELOAD_SECTION))
                elements = section.find_elements(*VIEW_LINKS)
                
                if not elements:
                    print(f"页面{page}无数据,结束遍历")
                    success = True
                    break

                # 提取URL
                for element in elements:
                    href = element.get_attribute('href')
                    if href:
                        href_set.add(href.replace('/Chome/', '/chome/'))

                # 追加写入文件(避免每次重写)
                with open('hrefs.txt', 'a', encoding='utf-8') as file:
                    existing_hrefs = set()
                    try:
                        with open('hrefs.txt', 'r', encoding='utf-8') as f:
                            existing_hrefs = set(f.read().splitlines())
                    except FileNotFoundError:
                        pass
                    new_hrefs = href_set - existing_hrefs
                    for href in new_hrefs:
                        file.write(href + '\n')

                # 跳转到下一页:优先用URL构造,避免依赖分页按钮
                page += 1
                next_page_url = NEXT_PAGE_TEMPLATE.format(page)
                driver.get(next_page_url)
                success = True
                break

            except (WebDriverException, TimeoutException, NoSuchElementException):
                print(f"页面{page}加载失败,重试第{attempt+1}次")
                driver.refresh()  # 刷新当前页重试

        if not success:
            print(f"页面{page}经过{max_retries}次重试仍失败,跳过")
            page += 1
            # 尝试直接跳转到下一页URL
            try:
                next_page_url = NEXT_PAGE_TEMPLATE.format(page)
                driver.get(next_page_url)
            except Exception as e:
                print(f"无法跳转到下一页,结束遍历: {str(e)}")
                break

finally:
    # 最后确保所有数据写入
    with open('hrefs.txt', 'w', encoding='utf-8') as file:
        for href in href_set:
            file.write(href + '\n')
    driver.quit()

代码说明

  • 显式等待:通过WebDriverWait等待元素加载,替代固定休眠,减少无效等待。
  • 无头模式:禁用UI渲染,大幅提升运行速度。
  • URL构造跳转:直接通过分页URL模板跳转,避免依赖分页按钮定位失败的问题。
  • 追加写入:仅写入新增的URL,减少IO操作次数。
  • 重试逻辑:失败后刷新当前页重试,3次失败后直接跳转到下一页URL,确保遍历不中断。

内容的提问来源于stack exchange,提问作者Tosh_gitonga

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 02:06:03