You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium+BeautifulSoup爬取网页耗时过长的优化咨询

问题场景

使用Selenium爬取Catch网站卖家页面的产品数据,流程为:

  1. 爬取列表页的产品链接存入hrefs数组
  2. 逐个访问产品链接,提取标题、价格、图片链接存入products数组

原代码如下:

爬取产品链接代码

from bs4 import BeautifulSoup
from selenium import webdriver
import os
chrome_options = webdriver.ChromeOptions()
chrome_options.add_argument("--headless")
service = webdriver.chrome.service.Service(executable_path=os.getcwd() + "./chromedriver.exe")
driver = webdriver.Chrome(service=service, options=chrome_options)
driver.set_page_load_timeout(900)
link = 'https://www.catch.com.au/seller/vdoo/products.html?page=1'
driver.get(link)
soup = BeautifulSoup(driver.page_source, 'lxml')
product_links = soup.find_all("a", class_="css-1k3ukvl")

hrefs = []
for product_link in product_links:
    href = product_link.get("href")
    if href.startswith("/"):
        href = "https://www.catch.com.au" + href
    hrefs.append(href)

爬取产品详情代码

products = []
for href in hrefs:
    driver.get(href)
    soup = BeautifulSoup(driver.page_source, 'lxml')
    
    title = soup.find("h1", class_="e12cshkt0").text.strip()
    price = soup.find("span", class_="css-1qfcjyj").text.strip()
    image_link = soup.find("img", class_="css-qvzl9f")["src"]
    product = {
        "title": title,
        "price": price,
        "image_link": image_link
    }
    products.append(product)
driver.quit()
print(len(products))

遇到的问题:

  • 单页爬取已耗时极长,甚至触发900秒超时
  • 扩展至40页列表+1440个产品页时,超时问题更严重
  • 逐个访问产品页的同步流程效率极低

优化方案

1. 用requests替代Selenium爬取静态内容

Selenium会加载页面所有资源(JS、图片、广告),而大部分数据是静态的,直接用requests+BeautifulSoup可以大幅减少加载时间。

优化后的链接爬取代码

import requests
from bs4 import BeautifulSoup

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

hrefs = []
# 爬取全部40页
for page in range(1, 41):
    url = f"https://www.catch.com.au/seller/vdoo/products.html?page={page}"
    response = requests.get(url, headers=headers)
    soup = BeautifulSoup(response.text, 'lxml')
    product_links = soup.find_all("a", class_="css-1k3ukvl")
    
    for link in product_links:
        href = link.get("href")
        if href.startswith("/"):
            href = "https://www.catch.com.au" + href
        hrefs.append(href)

# 保存链接到文件,方便后续分步处理
with open("product_links.txt", "w") as f:
    f.write("\n".join(hrefs))

优化后的产品详情爬取代码(同步版)

import requests
from bs4 import BeautifulSoup
import json

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

# 读取之前保存的链接
with open("product_links.txt", "r") as f:
    hrefs = [line.strip() for line in f if line.strip()]

products = []
for idx, href in enumerate(hrefs):
    try:
        response = requests.get(href, headers=headers)
        soup = BeautifulSoup(response.text, 'lxml')
        
        title = soup.find("h1", class_="e12cshkt0").text.strip()
        price = soup.find("span", class_="css-1qfcjyj").text.strip()
        image_link = soup.find("img", class_="css-qvzl9f")["src"]
        
        products.append({
            "title": title,
            "price": price,
            "image_link": image_link
        })
        print(f"已完成 {idx+1}/{len(hrefs)}")
    except Exception as e:
        print(f"爬取 {href} 失败: {str(e)}")

# 保存结果到JSON
with open("products.json", "w") as f:
    json.dump(products, f, indent=2)

2. 异步请求批量爬取产品页

用aiohttp实现异步请求,同时处理多个产品页,效率比同步提升数倍。

异步爬取代码

import aiohttp
import asyncio
from bs4 import BeautifulSoup
import json

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

async def fetch_product(session, href):
    try:
        async with session.get(href, headers=headers) as response:
            html = await response.text()
            soup = BeautifulSoup(html, 'lxml')
            
            title = soup.find("h1", class_="e12cshkt0").text.strip()
            price = soup.find("span", class_="css-1qfcjyj").text.strip()
            image_link = soup.find("img", class_="css-qvzl9f")["src"]
            
            return {
                "title": title,
                "price": price,
                "image_link": image_link
            }
    except Exception as e:
        print(f"爬取 {href} 失败: {str(e)}")
        return None

async def main():
    # 读取链接
    with open("product_links.txt", "r") as f:
        hrefs = [line.strip() for line in f if line.strip()]
    
    async with aiohttp.ClientSession() as session:
        tasks = [fetch_product(session, href) for href in hrefs]
        products = await asyncio.gather(*tasks)
        # 过滤掉爬取失败的项
        products = [p for p in products if p is not None]
    
    # 保存结果
    with open("products_async.json", "w") as f:
        json.dump(products, f, indent=2)
    print(f"共爬取 {len(products)} 个产品")

if __name__ == "__main__":
    asyncio.run(main())

3. 拆分程序为独立步骤

将流程拆分为两个完全独立的脚本:

  1. 链接爬取脚本:专门负责爬取所有产品链接并保存到文件(如product_links.txt)
  2. 数据爬取脚本:读取保存的链接,批量爬取产品详情并保存结果

这样做的好处:

  • 可以分开执行,避免某一步出错导致全部流程重跑
  • 链接爬取完成后,可以多次复用链接文件测试数据爬取逻辑
  • 便于分步优化和调试

4. 若必须使用Selenium的优化策略

如果网站内容是动态渲染(需要JS加载),必须用Selenium,可做以下优化:

  • 禁用图片加载:chrome_options.add_argument("--blink-settings=imagesEnabled=false")
  • 使用eager加载策略,只加载DOM结构,不等待全部资源:
    chrome_options.page_load_strategy = 'eager'
    
  • 限制加载超时时间,跳过无响应页面
  • 使用多线程/多进程并行处理产品页请求

内容的提问来源于stack exchange,提问作者Hashir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 00:00:09