如何利用Multithreading优化网页链接爬取代码提升执行效率
Hey there! Let's fix that slow function together. The main issue with your current code is that it makes sequential HTTP requests—each request has to wait for the previous one to finish before starting. Since web requests are mostly waiting for the server to respond (I/O bound work), multithreading is perfect here to overlap that wait time and cut down your total runtime drastically.
First, a quick correction: Your regex uses < (the HTML entity for <), but when you fetch the page with requests.get(), the response text contains actual < characters. I've fixed that in the code below—otherwise your initial link matching might not work correctly.
Step-by-Step Optimized Code
We'll use Python's built-in concurrent.futures.ThreadPoolExecutor to handle parallel requests, and refactor repetitive logic into a helper function for clarity.
import requests import re from concurrent.futures import ThreadPoolExecutor def get_pagination_links(url): """Helper function to fetch pagination links from a manufacturer page""" try: # Add a user-agent to avoid being blocked as a bot headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } req = requests.get(url, headers=headers) req.raise_for_status() # Raise error for HTTP status codes like 404/500 # Extract pagination section res2 = re.findall(r"<strong>1</strong>(.+)</div>", req.text) if not res2: return [] # Get all page links pages = re.findall(r"<a href=\"(.+?)\">.+</a>", res2[0]) return [f"https://www.gsmarena.com/{page}" for page in pages] except Exception as e: print(f"Failed to process {url}: {str(e)}") return [] def get_manufacturer(): # Fetch initial manufacturer links from the homepage homepage_url = "https://www.gsmarena.com/" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } manufacturers = requests.get(homepage_url, headers=headers) manufacturers.raise_for_status() # Extract manufacturer links (fixed regex to use actual < characters) res = re.findall(r"<li><a href=\"samsung-phones-9.php\">.+\n", manufacturers.text) manufacturer_links = re.findall(r"<li><a href=\"(.+?)\">", res[0]) final_list = [f"https://www.gsmarena.com/{link}" for link in manufacturer_links] # Use multithreading to fetch pagination links in parallel # Adjust max_workers based on your needs (don't set too high to avoid being blocked) with ThreadPoolExecutor(max_workers=10) as executor: # Map each manufacturer URL to a future task future_to_url = {executor.submit(get_pagination_links, url): url for url in final_list} # Collect results as they complete for future in ThreadPoolExecutor.as_completed(future_to_url): url = future_to_url[future] try: pagination_links = future.result() final_list.extend(pagination_links) except Exception as e: print(f"Error handling {url}: {str(e)}") return final_list
Key Improvements Explained
- Parallel Requests: The
ThreadPoolExecutorruns up to 10 requests at the same time. This means instead of waiting 20 seconds per request (hypothetically), you process 10 requests in that same 20 seconds. - Error Handling: Added
raise_for_status()to catch HTTP errors, and wrapped requests in try/except blocks to prevent the whole function from crashing if one request fails. - User-Agent Header: Added a browser-like user-agent to reduce the chance of being blocked by the website's anti-scraping measures.
Pro Tip: Ditch Regex for HTML Parsing
Regex is fragile for HTML—small changes to the website's structure will break your code. Consider using BeautifulSoup instead for more robust parsing. Here's how the get_pagination_links function would look with BeautifulSoup:
from bs4 import BeautifulSoup def get_pagination_links(url): try: headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } req = requests.get(url, headers=headers) req.raise_for_status() soup = BeautifulSoup(req.text, "html.parser") pagination_container = soup.find("div", class_="nav-pages") if not pagination_container: return [] # Extract all page links page_links = pagination_container.find_all("a") return [f"https://www.gsmarena.com/{link['href']}" for link in page_links] except Exception as e: print(f"Failed to process {url}: {str(e)}") return []
This version is easier to read and maintain, and won't break if the site adds extra whitespace or minor HTML changes.
Things to Keep in Mind
- Don't Overdo It: Adjust
max_workerscarefully. Setting it too high (like 100) can overwhelm the website and get your IP blocked. Start with 10-15 and adjust based on performance. - Rate Limiting: If you still get blocked, add small delays between requests (or use a library like
tenacityfor retries with backoff). - Respect
robots.txt: Check the site'srobots.txtto make sure you're allowed to scrape these pages.
内容的提问来源于stack exchange,提问作者zalts1

