使用Concurrent Futures抓网页遇异常后不继续执行如何解决?
问题根源
executor.map 返回的迭代器会按提交顺序返回任务结果,迭代到对应任务时如果该任务抛出异常,异常会直接向上抛出。你将异常捕获逻辑放在了整个for循环外层,触发一次异常后循环直接终止,后续未迭代的任务结果就不会被处理。
修复方案
有两种常用修改方式,都可以实现异常不中断整体流程:
方案1:在任务函数内部捕获异常(更推荐)
将异常处理逻辑封装在每个子任务中,不管执行成功失败都返回对应结果,外层迭代不会遇到异常,自然不会中断:
import concurrent.futures from urllib.request import urlopen from bs4 import BeautifulSoup CONNECTIONS = 8 archive_url_list = ["https://www.example.com", "https://www.example.com", "sdfihaslkhasd", "https://www.example.com"] archive_h1_list = [] def get_archive_h1(h1_url): try: html = urlopen(h1_url) bsh = BeautifulSoup(html.read(), 'lxml') return bsh.h1.text.strip() except Exception: return "Exception Error!" def concurrent_calls(): with concurrent.futures.ThreadPoolExecutor(max_workers=CONNECTIONS) as executor: for result in executor.map(get_archive_h1, archive_url_list): archive_h1_list.append(result) # 执行调用 concurrent_calls() print(archive_h1_list)
执行后输出就是你预期的 ['Example Domain', 'Example Domain', 'Exception Error!', 'Example Domain']。
方案2:将异常捕获放在迭代逻辑内部
如果不想修改原任务函数,可以把try-except移到迭代逻辑内部,每次取结果时单独捕获异常:
def concurrent_calls(): with concurrent.futures.ThreadPoolExecutor(max_workers=CONNECTIONS) as executor: f1 = executor.map(get_archive_h1, archive_url_list) while True: try: result = next(f1) archive_h1_list.append(result) except StopIteration: break except Exception: archive_h1_list.append("Exception Error!")
内容的提问来源于stack exchange,提问作者Lee Roy
相关产品推荐
相关产品推荐

