You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用Multithreading优化网页链接爬取代码提升执行效率

Boosting Your Web Scraping Speed with Multithreading

Hey there! Let's fix that slow function together. The main issue with your current code is that it makes sequential HTTP requests—each request has to wait for the previous one to finish before starting. Since web requests are mostly waiting for the server to respond (I/O bound work), multithreading is perfect here to overlap that wait time and cut down your total runtime drastically.

First, a quick correction: Your regex uses &lt; (the HTML entity for <), but when you fetch the page with requests.get(), the response text contains actual < characters. I've fixed that in the code below—otherwise your initial link matching might not work correctly.

Step-by-Step Optimized Code

We'll use Python's built-in concurrent.futures.ThreadPoolExecutor to handle parallel requests, and refactor repetitive logic into a helper function for clarity.

import requests
import re
from concurrent.futures import ThreadPoolExecutor

def get_pagination_links(url):
    """Helper function to fetch pagination links from a manufacturer page"""
    try:
        # Add a user-agent to avoid being blocked as a bot
        headers = {
            "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
        }
        req = requests.get(url, headers=headers)
        req.raise_for_status()  # Raise error for HTTP status codes like 404/500
        
        # Extract pagination section
        res2 = re.findall(r"<strong>1</strong>(.+)</div>", req.text)
        if not res2:
            return []
        
        # Get all page links
        pages = re.findall(r"<a href=\"(.+?)\">.+</a>", res2[0])
        return [f"https://www.gsmarena.com/{page}" for page in pages]
    
    except Exception as e:
        print(f"Failed to process {url}: {str(e)}")
        return []

def get_manufacturer():
    # Fetch initial manufacturer links from the homepage
    homepage_url = "https://www.gsmarena.com/"
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
    }
    
    manufacturers = requests.get(homepage_url, headers=headers)
    manufacturers.raise_for_status()
    
    # Extract manufacturer links (fixed regex to use actual < characters)
    res = re.findall(r"<li><a href=\"samsung-phones-9.php\">.+\n", manufacturers.text)
    manufacturer_links = re.findall(r"<li><a href=\"(.+?)\">", res[0])
    final_list = [f"https://www.gsmarena.com/{link}" for link in manufacturer_links]
    
    # Use multithreading to fetch pagination links in parallel
    # Adjust max_workers based on your needs (don't set too high to avoid being blocked)
    with ThreadPoolExecutor(max_workers=10) as executor:
        # Map each manufacturer URL to a future task
        future_to_url = {executor.submit(get_pagination_links, url): url for url in final_list}
        
        # Collect results as they complete
        for future in ThreadPoolExecutor.as_completed(future_to_url):
            url = future_to_url[future]
            try:
                pagination_links = future.result()
                final_list.extend(pagination_links)
            except Exception as e:
                print(f"Error handling {url}: {str(e)}")
    
    return final_list

Key Improvements Explained

  1. Parallel Requests: The ThreadPoolExecutor runs up to 10 requests at the same time. This means instead of waiting 20 seconds per request (hypothetically), you process 10 requests in that same 20 seconds.
  2. Error Handling: Added raise_for_status() to catch HTTP errors, and wrapped requests in try/except blocks to prevent the whole function from crashing if one request fails.
  3. User-Agent Header: Added a browser-like user-agent to reduce the chance of being blocked by the website's anti-scraping measures.

Pro Tip: Ditch Regex for HTML Parsing

Regex is fragile for HTML—small changes to the website's structure will break your code. Consider using BeautifulSoup instead for more robust parsing. Here's how the get_pagination_links function would look with BeautifulSoup:

from bs4 import BeautifulSoup

def get_pagination_links(url):
    try:
        headers = {
            "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
        }
        req = requests.get(url, headers=headers)
        req.raise_for_status()
        
        soup = BeautifulSoup(req.text, "html.parser")
        pagination_container = soup.find("div", class_="nav-pages")
        
        if not pagination_container:
            return []
        
        # Extract all page links
        page_links = pagination_container.find_all("a")
        return [f"https://www.gsmarena.com/{link['href']}" for link in page_links]
    
    except Exception as e:
        print(f"Failed to process {url}: {str(e)}")
        return []

This version is easier to read and maintain, and won't break if the site adds extra whitespace or minor HTML changes.

Things to Keep in Mind

  • Don't Overdo It: Adjust max_workers carefully. Setting it too high (like 100) can overwhelm the website and get your IP blocked. Start with 10-15 and adjust based on performance.
  • Rate Limiting: If you still get blocked, add small delays between requests (or use a library like tenacity for retries with backoff).
  • Respect robots.txt: Check the site's robots.txt to make sure you're allowed to scrape these pages.

内容的提问来源于stack exchange,提问作者zalts1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 06:47:24