You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在网页爬虫中运用多线程解决单线程爬虫耗时过长问题

Optimizing Your E-Commerce Crawler with Multithreading

Great question! Since web scraping is mostly I/O bound (spending most of its time waiting for HTTP responses), multithreading is perfect for speeding up your crawler. We can use Python's concurrent.futures.ThreadPoolExecutor to parallelize two key parts of your code:

  1. Running the Amazon and Flipkart scrapes in parallel instead of sequentially.
  2. Fetching Flipkart's individual product pages in parallel (the biggest bottleneck in your current code).

Here's how to modify your code step by step:

Step 1: Refactor into Reusable Functions

First, split your code into dedicated functions for scraping Amazon, scraping Flipkart product pages, and scraping individual Flipkart products. This makes it easy to run them in threads.

Modified Code

from Application import app
from flask import Flask, redirect, url_for, render_template, request, flash, jsonify
from bs4 import BeautifulSoup
import requests
import concurrent.futures  # Add this import

@app.route('/')
def root():
    return "Hello"

def scrape_amazon(search_string, headers):
    """Scrape product data from Amazon.in search results"""
    url = "https://www.amazon.in/s?k=" + search_string
    page = requests.get(url, headers=headers)
    soup = BeautifulSoup(page.text, "html.parser")
    results = soup.find_all('div', {'data-component-type': 's-search-result'})
    amazon_product = []
    
    for item in results:
        product = {}
        try:
            atag = item.h2.a
            description = atag.text.strip()
            rurl = 'https://www.amazon.in' + atag.get('href')
            
            price_parent = item.find('span', 'a-price')
            price = price_parent.find('span', 'a-offscreen').text
            
            rating = item.i.text
            review_count = item.find('span', {'class': 'a-size-base', 'dir': 'auto'}).text
        except AttributeError:
            continue  # Skip items with missing data
        
        product["description"] = description
        product["rurl"] = rurl
        product["price"] = price
        product["rating"] = rating
        product["review_count"] = review_count
        amazon_product.append(product)
    
    return amazon_product

def scrape_single_flipkart_product(product_url, headers):
    """Scrape data from a single Flipkart product page"""
    product = {}
    try:
        prod_page = requests.get(product_url, headers=headers)  # Added headers here!
        prod_soup = BeautifulSoup(prod_page.text, "html.parser")
        
        prod_des = prod_soup.find('span', {'class': "B_NuCI"}).text
        prod_price = prod_soup.find('div', {'class': "_30jeq3 _16Jk6d"}).text[1:]
        prod_rating = prod_soup.find('div', {'class': "_3LWZlK _3uSWvT"}).text
        prod_review = prod_soup.find('span', {'class': "_2_R_DZ"}).text
        
        product["description"] = prod_des
        product["rurl"] = product_url
        product["price"] = prod_price
        product["rating"] = prod_rating
        product["review_count"] = prod_review
    except AttributeError:
        # Handle missing fields gracefully
        pass
    except Exception as e:
        print(f"Error scraping {product_url}: {str(e)}")
    
    return product

def scrape_flipkart(search_string, headers):
    """Scrape product data from Flipkart.com, using threads for product pages"""
    url = "https://www.flipkart.com/search?q=" + search_string
    page = requests.get(url, headers=headers)
    soup = BeautifulSoup(page.text, "html.parser")
    results = soup.find_all('div', {"class": "_1AtVbE col-12-12"})
    product_name_links = []
    
    for i in results[3:-4]:
        a_tag = i.find('a')
        if a_tag:
            product_name_links.append("https://www.flipkart.com" + a_tag.get('href'))
    
    # Use ThreadPoolExecutor to fetch product pages in parallel
    flipkart_product = []
    # Limit threads to avoid triggering anti-scraping measures
    with concurrent.futures.ThreadPoolExecutor(max_workers=5) as executor:
        # Map the helper function to all product URLs
        results = executor.map(scrape_single_flipkart_product, product_name_links, [headers]*len(product_name_links))
        
        # Collect valid products
        for result in results:
            if result:
                flipkart_product.append(result)
    
    return flipkart_product

@app.route('/<string>')
def scrape_products(string):
    headers = {"User-Agent":"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/83.0.4103.61 Safari/537.36"}
    
    # Run Amazon and Flipkart scrapes in parallel
    with concurrent.futures.ThreadPoolExecutor(max_workers=2) as executor:
        # Submit both scrape tasks to the executor
        amazon_future = executor.submit(scrape_amazon, string, headers)
        flipkart_future = executor.submit(scrape_flipkart, string, headers)
        
        # Wait for and retrieve results
        amazon_data = amazon_future.result()
        flipkart_data = flipkart_future.result()
    
    payload = {
        'amazon': amazon_data,
        'flipkart': flipkart_data
    }
    return jsonify(payload)

Key Improvements & Notes

  1. Parallelized Flipkart Product Pages: Instead of fetching each product page one after another, we use ThreadPoolExecutor to run up to 5 requests at once. This will drastically reduce the time spent on Flipkart scraping.
  2. Parallel Amazon + Flipkart Scrapes: The main scrape for Amazon and Flipkart now runs in parallel, so you don't have to wait for one to finish before starting the other.
  3. Fixed Missing Headers: Your original code didn't pass headers to Flipkart product page requests—this could have led to blocks. We've added headers to all requests.
  4. Better Error Handling: Added try-except blocks to handle individual product scrape failures without breaking the entire process.
  5. Thread Limits: We set max_workers to 5 for product pages and 2 for main scrapes. Don't increase this too much (e.g., >10) as it may trigger anti-scraping systems on Amazon/Flipkart.

Additional Tips for Optimization

  • Reuse Requests Sessions: Use requests.Session() to persist connections between requests, which can speed up HTTP calls:
    session = requests.Session()
    session.headers.update(headers)
    # Then use session.get(url) instead of requests.get(url, headers=headers)
    
  • Add Retries: For temporary network issues, add retries using libraries like tenacity or requests.adapters.HTTPAdapter.
  • Rate Limiting: If you start getting blocked, add small delays between requests (e.g., time.sleep(0.5) inside the product scrape function) to mimic human behavior.
  • Rotate User-Agents: If you face frequent blocks, rotate between different User-Agent strings to avoid detection.

内容的提问来源于stack exchange,提问作者Varun Palla

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 10:42:35