如何为网页爬虫实现多线程提速?Pool的适用场景及植入位置
Great question! Let’s break this down clearly:
First, yes, multiprocessing.Pool is a feasible option for speeding up your crawler—but since web scraping is an I/O-bound task (most time is spent waiting for network responses), thread-based parallelism is usually more efficient. Threads have lower overhead than full processes, making them better suited for tasks where you’re waiting on external resources like web servers. That said, we’ll cover both approaches starting with the Pool implementation you asked about.
Step-by-Step Implementation with multiprocessing.Pool
To integrate Pool into your code, you’ll need to refactor your scraping logic into a reusable function, then use Pool.map() to run that function in parallel across your list of numbers. Here’s how to modify your original code:
Modified Code
import requests from multiprocessing import Pool def scrape_number(number): try: # Keeping your original URL structure (note: Google's URLs don't work this way, but we'll stick with it) url = f'https://www.google.com/{str(number)}' response = requests.get(url) response.raise_for_status() # Catch HTTP errors (4xx/5xx) return response.text except Exception as e: # Gracefully handle failures instead of crashing the whole crawl return f"Error scraping {number}: {str(e)}" if __name__ == '__main__': numbers = [4,8,5,7,3,10] # Initialize a pool of processes (adjust the number based on your system/website limits) with Pool(processes=4) as pool: # Distribute the numbers list across the pool and collect results results = pool.map(scrape_number, numbers) # Write results to file (same as your original logic) with open('testing.txt', 'w') as outfile: outfile.write("\n".join(results)) print(results)
Key Changes Explained:
- Extracted scraping logic into
scrape_number: This function encapsulates the work for a single number, which is required forPool.map()to distribute tasks across processes. - Wrapped execution in
if __name__ == '__main__':: Critical for Windows systems (prevents infinite process spawning). - Replaced the for loop with
pool.map(): This runs thescrape_numberfunction in parallel for every number in your list, instead of sequentially. - Added error handling: Ensures one failed request doesn’t break the entire crawl.
Better Alternative: Thread-Based Parallelism (For I/O-Bound Tasks)
Since web requests spend most of their time waiting, threads are a more efficient choice. You can use concurrent.futures.ThreadPoolExecutor (part of Python’s standard library) for this:
Example with ThreadPoolExecutor
import requests from concurrent.futures import ThreadPoolExecutor def scrape_number(number): try: url = f'https://www.google.com/{str(number)}' response = requests.get(url) response.raise_for_status() return response.text except Exception as e: return f"Error scraping {number}: {str(e)}" if __name__ == '__main__': numbers = [4,8,5,7,3,10] # Use a thread pool (adjust max_workers based on website tolerance) with ThreadPoolExecutor(max_workers=4) as executor: results = list(executor.map(scrape_number, numbers)) with open('testing.txt', 'w') as outfile: outfile.write("\n".join(results)) print(results)
Why Threads Are Better Here:
- Lower memory/CPU overhead compared to full processes.
- Threads share the same memory space, so no extra inter-process communication (IPC) costs.
- Perfect for tasks where you’re waiting on network responses or disk I/O.
Important Tips for Web Crawling:
- Rate Limiting: Add a small delay (e.g.,
time.sleep(1)) in thescrape_numberfunction to avoid getting blocked by websites. - Session Reuse: Use
requests.Session()to reuse TCP connections for faster repeated requests (safe for threads if you don’t modify the session concurrently). - Concurrency Limits: Don’t set
processesormax_workerstoo high—most websites block excessive concurrent requests. Start with 4-8 and adjust based on results.
内容的提问来源于stack exchange,提问作者Mia Ella

