如何为动态URL列表实现网页爬虫线程的启动与终止?
动态管理网页爬虫线程:新增/移除URL时启停对应线程
问题描述
我正在编写一款小型网页爬虫,希望实现当URL列表中添加或移除URL时,对应启动或终止线程的功能。目前已编写的代码如下:
import concurrent.futures import time import random import requests class WebScraper: def __init__(self): self.session = requests.Session() def run(self, url: str): while True: response = self.do_request(url) if response.status_code != 200: continue data = self.scrape_data(response) ... time.sleep(500) def do_request(self, url): response = self.session.get(url) return response def scrape_data(self, response): # TODO: Implement your web scraping logic here return {} if __name__ == '__main__': URLS_TO_TEST = [ "http://books.toscrape.com/catalogue/category/books/travel_2/index.html", "http://books.toscrape.com/catalogue/category/books/mystery_3/index.html", "http://books.toscrape.com/catalogue/category/books/historical-fiction_4/index.html", "http://books.toscrape.com/catalogue/category/books/sequential-art_5/index.html", "http://books.toscrape.com/catalogue/category/books/classics_6/index.html", ] with concurrent.futures.ThreadPoolExecutor() as executor: for url in URLS_TO_TEST: session = WebScraper() future = executor.submit(session.run, url) time.sleep(random.randint(10, 20)) URLS_TO_TEST.pop(random.randint(0, len(URLS_TO_TEST) - 1)) # The removed url should also terminate the thread time.sleep(random.randint(10, 20)) URLS_TO_TEST.append('http://books.toscrape.com/catalogue/category/books/health_47/index.html') # The added url should also start a new thread
我的需求是:URL从列表移除时终止对应线程,新增URL时启动对应线程,不确定能否通过主线程实现这一点,使用threading模块是否可行?后续计划通过数据库动态维护URL列表。
解决方案
1. threading模块完全可以实现该需求
Python没有提供强制终止线程的安全方法(强制终止可能导致资源泄漏、数据不一致),因此需要采用协作式终止的思路:让线程定期检查一个终止标志,当标志位触发时自行退出循环,结束线程。
2. 核心改造步骤
(1)给WebScraper类添加终止标志
修改WebScraper,加入线程运行状态标记,让run方法能响应终止信号:
import threading import requests import time class WebScraper: def __init__(self): self.session = requests.Session() self.running = True # 终止标志位 self.thread = None # 存储当前线程实例 def run(self, url: str): while self.running: try: response = self.do_request(url) if response.status_code == 200: data = self.scrape_data(response) # 处理爬取到的数据 print(f"Successfully scraped {url}") # 每次循环后检查终止标志,再进入休眠 if not self.running: break time.sleep(500) except Exception as e: print(f"Error scraping {url}: {e}") if not self.running: break time.sleep(60) # 出错后短时间重试 print(f"Stopped scraping {url}") def do_request(self, url): return self.session.get(url, timeout=10) def scrape_data(self, response): # 实现你的爬取逻辑 return {"url": response.url, "status": response.status_code} def stop(self): # 触发终止标志 self.running = False
(2)维护URL与Scraper实例的映射
用字典存储当前运行的URL对应的WebScraper实例,方便新增/移除时快速定位:
if __name__ == '__main__': # 存储URL到Scraper实例的映射 url_scraper_map = {} # 初始URL列表 initial_urls = [ "http://books.toscrape.com/catalogue/category/books/travel_2/index.html", "http://books.toscrape.com/catalogue/category/books/mystery_3/index.html", "http://books.toscrape.com/catalogue/category/books/historical-fiction_4/index.html", "http://books.toscrape.com/catalogue/category/books/sequential-art_5/index.html", "http://books.toscrape.com/catalogue/category/books/classics_6/index.html", ] # 启动初始URL的爬线程 for url in initial_urls: scraper = WebScraper() thread = threading.Thread(target=scraper.run, args=(url,), daemon=True) scraper.thread = thread thread.start() url_scraper_map[url] = scraper # 模拟动态移除URL time.sleep(10) removed_url = random.choice(list(url_scraper_map.keys())) print(f"Removing URL: {removed_url}") url_scraper_map[removed_url].stop() del url_scraper_map[removed_url] # 模拟动态新增URL time.sleep(10) new_url = "http://books.toscrape.com/catalogue/category/books/health_47/index.html" print(f"Adding URL: {new_url}") new_scraper = WebScraper() new_thread = threading.Thread(target=new_scraper.run, args=(new_url,), daemon=True) new_scraper.thread = new_thread new_thread.start() url_scraper_map[new_url] = new_scraper # 主线程保持运行 try: while True: time.sleep(3600) except KeyboardInterrupt: # 退出时停止所有线程 for scraper in url_scraper_map.values(): scraper.stop()
(3)结合数据库的动态维护思路
后续接入数据库时,可以在主线程中定期轮询数据库(比如每30秒一次):
- 对比数据库中的URL列表和当前
url_scraper_map的键 - 对于数据库中存在但本地没有的URL:启动新线程,加入映射
- 对于本地存在但数据库中已移除的URL:调用对应Scraper的
stop()方法,从映射中删除
关键注意事项
- 协作式终止的可靠性:确保
run方法中的循环每次都会检查self.running标志,尤其是在time.sleep前,避免线程长时间休眠无法响应终止信号。 - 守护线程设置:将爬线程设为守护线程,避免主线程退出时还有线程残留。
- 异常处理:在爬取逻辑中加入异常捕获,防止单个线程崩溃影响整个程序。
内容的提问来源于stack exchange,提问作者PythonNewbie
相关产品推荐
相关产品推荐

