You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为动态URL列表实现网页爬虫线程的启动与终止?

动态管理网页爬虫线程:新增/移除URL时启停对应线程

问题描述

我正在编写一款小型网页爬虫,希望实现当URL列表中添加或移除URL时,对应启动或终止线程的功能。目前已编写的代码如下:

import concurrent.futures
import time
import random

import requests

class WebScraper:
    def __init__(self):
        self.session = requests.Session()

    def run(self, url: str):
        while True:
            response = self.do_request(url)
            if response.status_code != 200:
                continue

            data = self.scrape_data(response)
            ...

            time.sleep(500)

    def do_request(self, url):
        response = self.session.get(url)
        return response

    def scrape_data(self, response):
        # TODO: Implement your web scraping logic here
        return {}

if __name__ == '__main__':
    URLS_TO_TEST = [
        "http://books.toscrape.com/catalogue/category/books/travel_2/index.html",
        "http://books.toscrape.com/catalogue/category/books/mystery_3/index.html",
        "http://books.toscrape.com/catalogue/category/books/historical-fiction_4/index.html",
        "http://books.toscrape.com/catalogue/category/books/sequential-art_5/index.html",
        "http://books.toscrape.com/catalogue/category/books/classics_6/index.html",
    ]
    with concurrent.futures.ThreadPoolExecutor() as executor:
        for url in URLS_TO_TEST:
            session = WebScraper()
            future = executor.submit(session.run, url)

    time.sleep(random.randint(10, 20))

    URLS_TO_TEST.pop(random.randint(0, len(URLS_TO_TEST) - 1))  # The removed url should also terminate the thread

    time.sleep(random.randint(10, 20))

    URLS_TO_TEST.append('http://books.toscrape.com/catalogue/category/books/health_47/index.html')  # The added url should also start a new thread

我的需求是:URL从列表移除时终止对应线程,新增URL时启动对应线程,不确定能否通过主线程实现这一点,使用threading模块是否可行?后续计划通过数据库动态维护URL列表。

解决方案

1. threading模块完全可以实现该需求

Python没有提供强制终止线程的安全方法(强制终止可能导致资源泄漏、数据不一致),因此需要采用协作式终止的思路:让线程定期检查一个终止标志,当标志位触发时自行退出循环,结束线程。

2. 核心改造步骤

(1)给WebScraper类添加终止标志

修改WebScraper,加入线程运行状态标记,让run方法能响应终止信号:

import threading
import requests
import time

class WebScraper:
    def __init__(self):
        self.session = requests.Session()
        self.running = True  # 终止标志位
        self.thread = None   # 存储当前线程实例

    def run(self, url: str):
        while self.running:
            try:
                response = self.do_request(url)
                if response.status_code == 200:
                    data = self.scrape_data(response)
                    # 处理爬取到的数据
                    print(f"Successfully scraped {url}")
                
                # 每次循环后检查终止标志,再进入休眠
                if not self.running:
                    break
                time.sleep(500)
            except Exception as e:
                print(f"Error scraping {url}: {e}")
                if not self.running:
                    break
                time.sleep(60)  # 出错后短时间重试
        print(f"Stopped scraping {url}")

    def do_request(self, url):
        return self.session.get(url, timeout=10)

    def scrape_data(self, response):
        # 实现你的爬取逻辑
        return {"url": response.url, "status": response.status_code}

    def stop(self):
        # 触发终止标志
        self.running = False

(2)维护URL与Scraper实例的映射

用字典存储当前运行的URL对应的WebScraper实例,方便新增/移除时快速定位:

if __name__ == '__main__':
    # 存储URL到Scraper实例的映射
    url_scraper_map = {}
    # 初始URL列表
    initial_urls = [
        "http://books.toscrape.com/catalogue/category/books/travel_2/index.html",
        "http://books.toscrape.com/catalogue/category/books/mystery_3/index.html",
        "http://books.toscrape.com/catalogue/category/books/historical-fiction_4/index.html",
        "http://books.toscrape.com/catalogue/category/books/sequential-art_5/index.html",
        "http://books.toscrape.com/catalogue/category/books/classics_6/index.html",
    ]

    # 启动初始URL的爬线程
    for url in initial_urls:
        scraper = WebScraper()
        thread = threading.Thread(target=scraper.run, args=(url,), daemon=True)
        scraper.thread = thread
        thread.start()
        url_scraper_map[url] = scraper

    # 模拟动态移除URL
    time.sleep(10)
    removed_url = random.choice(list(url_scraper_map.keys()))
    print(f"Removing URL: {removed_url}")
    url_scraper_map[removed_url].stop()
    del url_scraper_map[removed_url]

    # 模拟动态新增URL
    time.sleep(10)
    new_url = "http://books.toscrape.com/catalogue/category/books/health_47/index.html"
    print(f"Adding URL: {new_url}")
    new_scraper = WebScraper()
    new_thread = threading.Thread(target=new_scraper.run, args=(new_url,), daemon=True)
    new_scraper.thread = new_thread
    new_thread.start()
    url_scraper_map[new_url] = new_scraper

    # 主线程保持运行
    try:
        while True:
            time.sleep(3600)
    except KeyboardInterrupt:
        # 退出时停止所有线程
        for scraper in url_scraper_map.values():
            scraper.stop()

(3)结合数据库的动态维护思路

后续接入数据库时,可以在主线程中定期轮询数据库(比如每30秒一次):

  • 对比数据库中的URL列表和当前url_scraper_map的键
  • 对于数据库中存在但本地没有的URL:启动新线程,加入映射
  • 对于本地存在但数据库中已移除的URL:调用对应Scraper的stop()方法,从映射中删除

关键注意事项

  • 协作式终止的可靠性:确保run方法中的循环每次都会检查self.running标志,尤其是在time.sleep前,避免线程长时间休眠无法响应终止信号。
  • 守护线程设置:将爬线程设为守护线程,避免主线程退出时还有线程残留。
  • 异常处理:在爬取逻辑中加入异常捕获,防止单个线程崩溃影响整个程序。

内容的提问来源于stack exchange,提问作者PythonNewbie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 21:38:11