如何在网页爬虫中运用多线程解决单线程爬虫耗时过长问题
Optimizing Your E-Commerce Crawler with Multithreading
Great question! Since web scraping is mostly I/O bound (spending most of its time waiting for HTTP responses), multithreading is perfect for speeding up your crawler. We can use Python's concurrent.futures.ThreadPoolExecutor to parallelize two key parts of your code:
- Running the Amazon and Flipkart scrapes in parallel instead of sequentially.
- Fetching Flipkart's individual product pages in parallel (the biggest bottleneck in your current code).
Here's how to modify your code step by step:
Step 1: Refactor into Reusable Functions
First, split your code into dedicated functions for scraping Amazon, scraping Flipkart product pages, and scraping individual Flipkart products. This makes it easy to run them in threads.
Modified Code
from Application import app from flask import Flask, redirect, url_for, render_template, request, flash, jsonify from bs4 import BeautifulSoup import requests import concurrent.futures # Add this import @app.route('/') def root(): return "Hello" def scrape_amazon(search_string, headers): """Scrape product data from Amazon.in search results""" url = "https://www.amazon.in/s?k=" + search_string page = requests.get(url, headers=headers) soup = BeautifulSoup(page.text, "html.parser") results = soup.find_all('div', {'data-component-type': 's-search-result'}) amazon_product = [] for item in results: product = {} try: atag = item.h2.a description = atag.text.strip() rurl = 'https://www.amazon.in' + atag.get('href') price_parent = item.find('span', 'a-price') price = price_parent.find('span', 'a-offscreen').text rating = item.i.text review_count = item.find('span', {'class': 'a-size-base', 'dir': 'auto'}).text except AttributeError: continue # Skip items with missing data product["description"] = description product["rurl"] = rurl product["price"] = price product["rating"] = rating product["review_count"] = review_count amazon_product.append(product) return amazon_product def scrape_single_flipkart_product(product_url, headers): """Scrape data from a single Flipkart product page""" product = {} try: prod_page = requests.get(product_url, headers=headers) # Added headers here! prod_soup = BeautifulSoup(prod_page.text, "html.parser") prod_des = prod_soup.find('span', {'class': "B_NuCI"}).text prod_price = prod_soup.find('div', {'class': "_30jeq3 _16Jk6d"}).text[1:] prod_rating = prod_soup.find('div', {'class': "_3LWZlK _3uSWvT"}).text prod_review = prod_soup.find('span', {'class': "_2_R_DZ"}).text product["description"] = prod_des product["rurl"] = product_url product["price"] = prod_price product["rating"] = prod_rating product["review_count"] = prod_review except AttributeError: # Handle missing fields gracefully pass except Exception as e: print(f"Error scraping {product_url}: {str(e)}") return product def scrape_flipkart(search_string, headers): """Scrape product data from Flipkart.com, using threads for product pages""" url = "https://www.flipkart.com/search?q=" + search_string page = requests.get(url, headers=headers) soup = BeautifulSoup(page.text, "html.parser") results = soup.find_all('div', {"class": "_1AtVbE col-12-12"}) product_name_links = [] for i in results[3:-4]: a_tag = i.find('a') if a_tag: product_name_links.append("https://www.flipkart.com" + a_tag.get('href')) # Use ThreadPoolExecutor to fetch product pages in parallel flipkart_product = [] # Limit threads to avoid triggering anti-scraping measures with concurrent.futures.ThreadPoolExecutor(max_workers=5) as executor: # Map the helper function to all product URLs results = executor.map(scrape_single_flipkart_product, product_name_links, [headers]*len(product_name_links)) # Collect valid products for result in results: if result: flipkart_product.append(result) return flipkart_product @app.route('/<string>') def scrape_products(string): headers = {"User-Agent":"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/83.0.4103.61 Safari/537.36"} # Run Amazon and Flipkart scrapes in parallel with concurrent.futures.ThreadPoolExecutor(max_workers=2) as executor: # Submit both scrape tasks to the executor amazon_future = executor.submit(scrape_amazon, string, headers) flipkart_future = executor.submit(scrape_flipkart, string, headers) # Wait for and retrieve results amazon_data = amazon_future.result() flipkart_data = flipkart_future.result() payload = { 'amazon': amazon_data, 'flipkart': flipkart_data } return jsonify(payload)
Key Improvements & Notes
- Parallelized Flipkart Product Pages: Instead of fetching each product page one after another, we use
ThreadPoolExecutorto run up to 5 requests at once. This will drastically reduce the time spent on Flipkart scraping. - Parallel Amazon + Flipkart Scrapes: The main scrape for Amazon and Flipkart now runs in parallel, so you don't have to wait for one to finish before starting the other.
- Fixed Missing Headers: Your original code didn't pass headers to Flipkart product page requests—this could have led to blocks. We've added headers to all requests.
- Better Error Handling: Added try-except blocks to handle individual product scrape failures without breaking the entire process.
- Thread Limits: We set
max_workersto 5 for product pages and 2 for main scrapes. Don't increase this too much (e.g., >10) as it may trigger anti-scraping systems on Amazon/Flipkart.
Additional Tips for Optimization
- Reuse Requests Sessions: Use
requests.Session()to persist connections between requests, which can speed up HTTP calls:session = requests.Session() session.headers.update(headers) # Then use session.get(url) instead of requests.get(url, headers=headers) - Add Retries: For temporary network issues, add retries using libraries like
tenacityorrequests.adapters.HTTPAdapter. - Rate Limiting: If you start getting blocked, add small delays between requests (e.g.,
time.sleep(0.5)inside the product scrape function) to mimic human behavior. - Rotate User-Agents: If you face frequent blocks, rotate between different User-Agent strings to avoid detection.
内容的提问来源于stack exchange,提问作者Varun Palla
相关产品推荐
相关产品推荐

