使用Selenium+BeautifulSoup爬取网页耗时过长的优化咨询
问题场景
使用Selenium爬取Catch网站卖家页面的产品数据,流程为:
- 爬取列表页的产品链接存入
hrefs数组 - 逐个访问产品链接,提取标题、价格、图片链接存入
products数组
原代码如下:
爬取产品链接代码
from bs4 import BeautifulSoup from selenium import webdriver import os chrome_options = webdriver.ChromeOptions() chrome_options.add_argument("--headless") service = webdriver.chrome.service.Service(executable_path=os.getcwd() + "./chromedriver.exe") driver = webdriver.Chrome(service=service, options=chrome_options) driver.set_page_load_timeout(900) link = 'https://www.catch.com.au/seller/vdoo/products.html?page=1' driver.get(link) soup = BeautifulSoup(driver.page_source, 'lxml') product_links = soup.find_all("a", class_="css-1k3ukvl") hrefs = [] for product_link in product_links: href = product_link.get("href") if href.startswith("/"): href = "https://www.catch.com.au" + href hrefs.append(href)
爬取产品详情代码
products = [] for href in hrefs: driver.get(href) soup = BeautifulSoup(driver.page_source, 'lxml') title = soup.find("h1", class_="e12cshkt0").text.strip() price = soup.find("span", class_="css-1qfcjyj").text.strip() image_link = soup.find("img", class_="css-qvzl9f")["src"] product = { "title": title, "price": price, "image_link": image_link } products.append(product) driver.quit() print(len(products))
遇到的问题:
- 单页爬取已耗时极长,甚至触发900秒超时
- 扩展至40页列表+1440个产品页时,超时问题更严重
- 逐个访问产品页的同步流程效率极低
优化方案
1. 用requests替代Selenium爬取静态内容
Selenium会加载页面所有资源(JS、图片、广告),而大部分数据是静态的,直接用requests+BeautifulSoup可以大幅减少加载时间。
优化后的链接爬取代码
import requests from bs4 import BeautifulSoup headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } hrefs = [] # 爬取全部40页 for page in range(1, 41): url = f"https://www.catch.com.au/seller/vdoo/products.html?page={page}" response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'lxml') product_links = soup.find_all("a", class_="css-1k3ukvl") for link in product_links: href = link.get("href") if href.startswith("/"): href = "https://www.catch.com.au" + href hrefs.append(href) # 保存链接到文件,方便后续分步处理 with open("product_links.txt", "w") as f: f.write("\n".join(hrefs))
优化后的产品详情爬取代码(同步版)
import requests from bs4 import BeautifulSoup import json headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } # 读取之前保存的链接 with open("product_links.txt", "r") as f: hrefs = [line.strip() for line in f if line.strip()] products = [] for idx, href in enumerate(hrefs): try: response = requests.get(href, headers=headers) soup = BeautifulSoup(response.text, 'lxml') title = soup.find("h1", class_="e12cshkt0").text.strip() price = soup.find("span", class_="css-1qfcjyj").text.strip() image_link = soup.find("img", class_="css-qvzl9f")["src"] products.append({ "title": title, "price": price, "image_link": image_link }) print(f"已完成 {idx+1}/{len(hrefs)}") except Exception as e: print(f"爬取 {href} 失败: {str(e)}") # 保存结果到JSON with open("products.json", "w") as f: json.dump(products, f, indent=2)
2. 异步请求批量爬取产品页
用aiohttp实现异步请求,同时处理多个产品页,效率比同步提升数倍。
异步爬取代码
import aiohttp import asyncio from bs4 import BeautifulSoup import json headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } async def fetch_product(session, href): try: async with session.get(href, headers=headers) as response: html = await response.text() soup = BeautifulSoup(html, 'lxml') title = soup.find("h1", class_="e12cshkt0").text.strip() price = soup.find("span", class_="css-1qfcjyj").text.strip() image_link = soup.find("img", class_="css-qvzl9f")["src"] return { "title": title, "price": price, "image_link": image_link } except Exception as e: print(f"爬取 {href} 失败: {str(e)}") return None async def main(): # 读取链接 with open("product_links.txt", "r") as f: hrefs = [line.strip() for line in f if line.strip()] async with aiohttp.ClientSession() as session: tasks = [fetch_product(session, href) for href in hrefs] products = await asyncio.gather(*tasks) # 过滤掉爬取失败的项 products = [p for p in products if p is not None] # 保存结果 with open("products_async.json", "w") as f: json.dump(products, f, indent=2) print(f"共爬取 {len(products)} 个产品") if __name__ == "__main__": asyncio.run(main())
3. 拆分程序为独立步骤
将流程拆分为两个完全独立的脚本:
- 链接爬取脚本:专门负责爬取所有产品链接并保存到文件(如
product_links.txt) - 数据爬取脚本:读取保存的链接,批量爬取产品详情并保存结果
这样做的好处:
- 可以分开执行,避免某一步出错导致全部流程重跑
- 链接爬取完成后,可以多次复用链接文件测试数据爬取逻辑
- 便于分步优化和调试
4. 若必须使用Selenium的优化策略
如果网站内容是动态渲染(需要JS加载),必须用Selenium,可做以下优化:
- 禁用图片加载:
chrome_options.add_argument("--blink-settings=imagesEnabled=false") - 使用
eager加载策略,只加载DOM结构,不等待全部资源:chrome_options.page_load_strategy = 'eager' - 限制加载超时时间,跳过无响应页面
- 使用多线程/多进程并行处理产品页请求
内容的提问来源于stack exchange,提问作者Hashir
相关产品推荐
相关产品推荐

