Python编写爬虫时函数内向线程池URL数组添加元素失效问题求解
问题原因和修复方案
核心问题点
- 线程池初始仅提交了从文件读取的初始URL任务,后续追加到
urls数组的新链接没有被重新提交给线程池,自然不会被处理 - 初始代码提交完初始URL后直接调用了
e.shutdown(),线程池进入关闭状态,不再接收新的任务提交,后续新解析到的链接根本无法提交给线程池处理 - 追加链接的逻辑错误:直接调用
urls.append(links)是将整个links列表作为单个元素追加到urls中,会导致urls变为嵌套数组,无法正常遍历使用 - 相对路径处理逻辑错误:
link.get('href')返回的是字符串,Python中字符串为不可变类型,无法通过link[0] = ""修改第一个字符,会直接抛出类型错误 - 多线程操作全局
urls数组没有加锁,高并发场景下会出现数据丢失、脏写问题 - 没有去重逻辑,同一个URL可能被反复提交爬取,造成资源浪费
修复后的可运行代码
from concurrent import futures from urllib.request import Request, urlopen from bs4 import BeautifulSoup import threading from urllib.parse import urljoin # 全局资源声明 urls = [] crawled = set() lock = threading.Lock() executor = None def linksSearchAndAppend(url): global urls, crawled, lock, executor # 线程安全校验,避免重复爬取 with lock: if url in crawled: return crawled.add(url) try: # 增加UA和超时处理,避免被反爬或无响应卡住 req = Request(url, headers={'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'}) html_page = urlopen(req, timeout=10) soup = BeautifulSoup(html_page, "lxml") except Exception as e: print(f"爬取URL {url} 失败:{str(e)}") return for link in soup.findAll('a'): href = link.get('href') # 过滤空href if not href: continue # 自动拼接绝对路径,比手动拼接容错率更高 absolute_url = urljoin(url, href) # 过滤非http/https的其他协议链接 if not absolute_url.startswith(('http://', 'https://')): continue # 线程安全新增待爬链接并提交任务 with lock: if absolute_url not in crawled and absolute_url not in urls: urls.append(absolute_url) executor.submit(linksSearchAndAppend, absolute_url) print(f"当前累计待爬URL数:{len(urls)},已完成爬取URL数:{len(crawled)}") if __name__ == "__main__": # 读取初始URL列表 with open("urlList.txt", "r", encoding="utf-8") as f: for line in f: url = line.rstrip() if url: urls.append(url) # 初始化线程池 executor = futures.ThreadPoolExecutor(max_workers=8) # 提交初始爬取任务 for url in urls: executor.submit(linksSearchAndAppend, url) # 等待所有爬取任务全部完成后再关闭线程池 executor.shutdown(wait=True) # 最终把所有采集到的URL写入文件保存 with open("all_crawled_urls.txt", "w", encoding="utf-8") as f: for url in urls: f.write(url + "\n")
关键改动说明
- 新增线程锁
threading.Lock,保证全局URL数组和已爬取集合的线程安全,避免并发场景下的数据异常 - 新增
crawled集合存储已经爬取过的URL,避免重复爬取浪费资源 - 用标准库
urljoin拼接绝对路径,替代原有的手动字符串修改逻辑,兼容性和容错率更高 - 每次解析到有效的新URL后,直接提交给线程池执行任务,同时追加到全局URL数组
- 增加请求UA、超时、异常捕获逻辑,避免爬虫被反爬拦截或者无响应卡住
- 调整线程池关闭时机,等所有爬取任务全部执行完成后再调用
shutdown关闭线程池 - 增加最终URL结果的持久化存储,方便后续使用
内容的提问来源于stack exchange,提问作者Maaxkar
相关产品推荐
相关产品推荐

