如何高效抓取单网站海量URL?Selenium优化及重试方案咨询
优化10000+页面的Selenium URL提取方案
现有Selenium代码可遍历网站分页提取URL,但面对10000+页面时耗时极长;需要实现页面点击失败3次后跳过该页继续下一页的逻辑,且网站无公开API。
关键优化措施
1. 用显式等待替代固定休眠
固定time.sleep()会浪费大量等待时间,改用WebDriverWait等待目标元素加载完成,仅在必要时等待,大幅提升遍历效率。
2. 启用无头浏览器模式
禁用Chrome的UI渲染,减少资源占用,运行速度至少提升30%以上。
3. 优化文件IO操作
原代码每次循环都重写整个文件,改为追加写入或定时批量写入,避免重复IO开销。
4. 修复分页重试逻辑
原代码重试失败后会终止循环,调整为:3次重试失败后,直接尝试跳转到下一页(优先通过URL构造,避免依赖分页按钮),确保遍历不中断。
5. 精简元素定位逻辑
复用定位表达式,避免重复编写,同时使用更精准的定位方式(比如相对XPath)减少元素查找失败概率。
改进后的完整代码
import random from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from selenium.common.exceptions import WebDriverException, NoSuchElementException, TimeoutException user_agents = [ "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/109.0.0.0 Safari/537.36", "Mozilla/5.0 (Macintosh; Intel Mac OS X 13_1) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/16.1 Safari/605.1.15" ] # 配置Chrome选项 chrome_options = Options() chrome_options.add_argument(f'user-agent={random.choice(user_agents)}') chrome_options.add_argument('--headless=new') # 无头模式 chrome_options.add_argument('--disable-gpu') chrome_options.add_argument('--no-sandbox') chrome_options.add_argument('--disable-dev-shm-usage') driver = webdriver.Chrome(options=chrome_options) wait = WebDriverWait(driver, 10) # 显式等待超时时间 main_url = 'https://sekolah.data.kemdikbud.go.id/index.php/Chome/pencarian/' href_set = set() page = 1 max_retries = 3 # 提前定义定位表达式 PAGELOAD_SECTION = (By.ID, 'pageload') VIEW_LINKS = (By.XPATH, './/li[@class="list-group-item"]/a[contains(., "Lihat")]') # 相对路径定位 NEXT_PAGE_TEMPLATE = 'https://sekolah.data.kemdikbud.go.id/index.php/Chome/pencarian/{}' try: driver.get(main_url) # 等待搜索按钮并点击 search_button = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, 'button[type="submit"]'))) search_button.click() while True: success = False for attempt in range(max_retries): try: # 等待页面内容加载 section = wait.until(EC.presence_of_element_located(PAGELOAD_SECTION)) elements = section.find_elements(*VIEW_LINKS) if not elements: print(f"页面{page}无数据,结束遍历") success = True break # 提取URL for element in elements: href = element.get_attribute('href') if href: href_set.add(href.replace('/Chome/', '/chome/')) # 追加写入文件(避免每次重写) with open('hrefs.txt', 'a', encoding='utf-8') as file: existing_hrefs = set() try: with open('hrefs.txt', 'r', encoding='utf-8') as f: existing_hrefs = set(f.read().splitlines()) except FileNotFoundError: pass new_hrefs = href_set - existing_hrefs for href in new_hrefs: file.write(href + '\n') # 跳转到下一页:优先用URL构造,避免依赖分页按钮 page += 1 next_page_url = NEXT_PAGE_TEMPLATE.format(page) driver.get(next_page_url) success = True break except (WebDriverException, TimeoutException, NoSuchElementException): print(f"页面{page}加载失败,重试第{attempt+1}次") driver.refresh() # 刷新当前页重试 if not success: print(f"页面{page}经过{max_retries}次重试仍失败,跳过") page += 1 # 尝试直接跳转到下一页URL try: next_page_url = NEXT_PAGE_TEMPLATE.format(page) driver.get(next_page_url) except Exception as e: print(f"无法跳转到下一页,结束遍历: {str(e)}") break finally: # 最后确保所有数据写入 with open('hrefs.txt', 'w', encoding='utf-8') as file: for href in href_set: file.write(href + '\n') driver.quit()
代码说明
- 显式等待:通过
WebDriverWait等待元素加载,替代固定休眠,减少无效等待。 - 无头模式:禁用UI渲染,大幅提升运行速度。
- URL构造跳转:直接通过分页URL模板跳转,避免依赖分页按钮定位失败的问题。
- 追加写入:仅写入新增的URL,减少IO操作次数。
- 重试逻辑:失败后刷新当前页重试,3次失败后直接跳转到下一页URL,确保遍历不中断。
内容的提问来源于stack exchange,提问作者Tosh_gitonga
相关产品推荐
相关产品推荐

