You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C#多线程控制台爬虫改造技术咨询:单线程处理海量URL遇瓶颈

Hey there! As someone who’s tinkered with multi-threaded scrapers and dealt with exactly this kind of MySQL-backed URL queue, let me walk you through a straightforward, beginner-friendly solution to get your property crawler scaled up for hundreds of thousands of URLs.

Core Transformation Approach & Step-by-Step Implementation

1. Fix MySQL Concurrency Safety First

The biggest pitfall with multi-threaded crawlers here is duplicate URL processing. Multiple threads trying to fetch the same unprocessed URL at the same time will waste resources and create messy data. Here’s how to fix it with atomic database operations:

Use SELECT ... FOR UPDATE SKIP LOCKED to atomically grab unprocessed URLs while locking them so other threads can’t access them:

SELECT id, url FROM property_urls WHERE processed = 0 LIMIT 10 FOR UPDATE SKIP LOCKED;

After processing a URL, update its status safely:

UPDATE property_urls SET processed = 1, processed_at = NOW() WHERE id = ?;

For Older MySQL Versions (No SKIP LOCKED)

Mark URLs as "in progress" first, then fetch them:

-- Mark 10 URLs as being processed
UPDATE property_urls SET processed = 2 WHERE processed = 0 LIMIT 10;
-- Fetch the marked URLs
SELECT id, url FROM property_urls WHERE processed = 2;

After processing, set processed = 1 for success, or processed = 3 for failure (to retry later).

2. Use ThreadPoolExecutor for Easy Multi-Threading

Python’s concurrent.futures.ThreadPoolExecutor is perfect for beginners—it handles thread management so you don’t have to manually create Thread objects. Here’s a complete code framework tailored to your use case:

First, set up dependencies and a database connection pool (critical to avoid overwhelming MySQL with too many connections):

import mysql.connector
from mysql.connector.pooling import MySQLConnectionPool
from concurrent.futures import ThreadPoolExecutor
import requests
from bs4 import BeautifulSoup  # Replace with your parsing tool if needed
import time
import random

# Initialize connection pool (tweak pool size to match your thread count)
db_pool = MySQLConnectionPool(
    pool_name="property_crawler_pool",
    pool_size=15,
    host="your_db_host",
    user="your_db_user",
    password="your_db_password",
    database="your_db_name"
)

Next, write the function that handles a single URL’s scraping and database updates:

def process_single_url(url_id, url):
    try:
        # Grab a connection from the pool
        conn = db_pool.get_connection()
        cursor = conn.cursor()

        # 1. Scrape the property page (replace with your actual scraping logic)
        headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"}
        response = requests.get(url, headers=headers, timeout=15)
        response.raise_for_status()  # Trigger exception for HTTP errors
        
        soup = BeautifulSoup(response.text, "html.parser")
        # Parse your property data (example fields—adjust to your needs)
        property_data = {
            "url_id": url_id,
            "price": soup.find("span", class_="property-price").text.strip(),
            "area": soup.find("div", class_="property-area").text.strip()
        }

        # 2. Insert parsed data into your property table
        insert_query = """
        INSERT INTO property_data (url_id, price, area)
        VALUES (%(url_id)s, %(price)s, %(area)s)
        """
        cursor.execute(insert_query, property_data)

        # 3. Mark the URL as processed
        update_query = "UPDATE property_urls SET processed = 1 WHERE id = %s"
        cursor.execute(update_query, (url_id,))

        conn.commit()
        print(f"✅ Processed URL: {url}")

    except Exception as e:
        # Log failures and mark URL as failed (for later retries)
        print(f"❌ Failed to process {url}: {str(e)}")
        if "conn" in locals():
            cursor.execute("UPDATE property_urls SET processed = 3 WHERE id = %s", (url_id,))
            conn.commit()
    finally:
        # Always clean up connections and cursors to return to the pool
        if "cursor" in locals():
            cursor.close()
        if "conn" in locals():
            conn.close()
        # Add a small random delay to avoid triggering anti-scraping measures
        time.sleep(random.uniform(0.5, 2))

Finally, set up the thread pool to process URLs in batches:

def main():
    max_threads = 15  # Start small (5-10) and adjust based on server/network capacity
    batch_size = 10   # Number of URLs to fetch per batch

    with ThreadPoolExecutor(max_workers=max_threads) as executor:
        while True:
            # Fetch a batch of unprocessed URLs with atomic locking
            conn = db_pool.get_connection()
            cursor = conn.cursor(dictionary=True)
            cursor.execute("SELECT id, url FROM property_urls WHERE processed = 0 LIMIT %s FOR UPDATE SKIP LOCKED", (batch_size,))
            url_batch = cursor.fetchall()
            cursor.close()
            conn.close()

            if not url_batch:
                print("\nAll URLs processed! 🎉")
                break

            # Submit batch to thread pool
            futures = [executor.submit(process_single_url, item["id"], item["url"]) for item in url_batch]
            # Wait for current batch to finish before fetching next (optional but easier to monitor)
            for future in futures:
                future.result()

if __name__ == "__main__":
    main()

3. Critical Pitfalls to Avoid as a Beginner

  • Don’t skip connection pools: Creating a new database connection for every thread will crash your MySQL server. The pool reuses connections efficiently.
  • Don’t set thread count too high: More threads don’t equal faster processing—you’ll hit database/network limits or get blocked by the target website. Start with 5-10 and tune up slowly.
  • Always handle exceptions: Network errors, broken page structures, and timeouts are inevitable. Failing to catch them will crash threads and leave URLs in limbo.
  • Respect anti-scraping rules: Add user agents, random delays, and avoid hammering the target site. If you get blocked, consider adding a proxy pool or reducing thread count.

内容的提问来源于stack exchange,提问作者Subrato M

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:22:01