You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Selenium与多进程的Python网络爬虫运行异常求助

Hey there! Let's work through your multi-threaded crawler issues step by step:

1. Multiple Browser Windows Popping Up

First off, those 4 browser windows are totally expected if you're using a tool like Selenium with multi-threading. Each thread is spawning its own browser instance to handle a URL. If you want to avoid this (and save system resources), switch to headless mode for your browser:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

def init_browser():
    chrome_options = Options()
    # Enable headless mode (no visible window)
    chrome_options.add_argument("--headless=new")
    chrome_options.add_argument("--disable-gpu")
    chrome_options.add_argument("--no-sandbox")  # Useful for Linux environments
    return webdriver.Chrome(options=chrome_options)

You can also tweak the number of threads (e.g., in ThreadPoolExecutor) to control how many browser instances run at once—don't go too high, as it'll eat up RAM and CPU.

2. The "recor..." Error (Most Likely Thread Safety or Recursion Issues)

Since you only mentioned a partial error message, here are the two most common culprits:

Case 1: RecursionError

If your crawler function uses recursion (e.g., following links recursively), multi-threading can trigger this faster than single-threading. Each thread has its own recursion depth limit, and running multiple recursive calls at once can hit Python's default limit (~1000) quicker.

  • Fix: Check your code for accidental recursive calls. If recursion is intentional, increase the limit temporarily with sys.setrecursionlimit(10000) (but use this cautiously—it can cause crashes if overdone).

Case 2: Thread-Safety Issues with base_list

Python lists are not thread-safe. If multiple threads are appending data to base_list at the same time, you'll get corrupted data, missing entries, or errors related to record/state inconsistencies.

  • Fix: Use a threading.Lock to protect access to the list:
import threading

base_list = []
# Create a lock object
list_lock = threading.Lock()

def crawl(url):
    # Your existing code to scrape data from the URL
    scraped_data = your_scraping_logic(url)
    
    # Only modify the list while holding the lock
    with list_lock:
        base_list.append(scraped_data)

This ensures only one thread can write to base_list at a time, preventing race conditions.

Quick Pro Tip for Managing 1000 URLs

Use concurrent.futures.ThreadPoolExecutor to handle thread management cleanly—you can set a reasonable max worker count (start with 8-12, adjust based on your machine's specs):

from concurrent.futures import ThreadPoolExecutor

def main():
    url_list = [your 1000 URLs here]
    with ThreadPoolExecutor(max_workers=8) as executor:
        executor.map(crawl, url_list)
    
    # After all threads finish, base_list has all your data
    print(f"Scraped {len(base_list)} entries")

If you can share the full error traceback, I can narrow this down even more!

内容的提问来源于stack exchange,提问作者matthew matthee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:52:26