如何让Python脚本遇urllib.request HTTP错误时继续执行?
解决URL探测脚本遇HTTP错误终止的问题
我有个读取URL列表CSV的Python脚本,CSV内容如下:
name, url google, https://httpstat.us/200 yahoo, https://httpstat.us/401 bcs, https://httpstat.us/521
当前脚本运行时,遇到返回HTTP错误码的URL就直接终止,无法处理剩余URL,报错信息如下:
➜ DevOps_Practice python urlProbe.py 1687386788.8879528 Counter: 1 Website: google URL: https://httpstat.us/200 200 200 All is good! google - 1 - 0 - 0 - 0 Website: yahoo URL: https://httpstat.us/401 Traceback (most recent call last): File "/Users/desmondlim/Documents/DevOps/Projects/DevOps_Challenge_SPH/urlProbe.py", line 38, in <module> print(urllib.request.urlopen(url_entry["url"]).status) File "/Users/desmondlim/.pyenv/versions/3.10.4/lib/python3.10/urllib/request.py", line 216, in urlopen return opener.open(url, data, timeout) File "/Users/desmondlim/.pyenv/versions/3.10.4/lib/python3.10/urllib/request.py", line 525, in open response = meth(req, response) File "/Users/desmondlim/.pyenv/versions/3.10.4/lib/python3.10/urllib/request.py", line 634, in http_response response = self.parent.error( File "/Users/desmondlim/.pyenv/versions/3.10.4/lib/python3.10/urllib/request.py", line 563, in error return self._call_chain(*args) File "/Users/desmondlim/.pyenv/versions/3.10.4/lib/python3.10/urllib/request.py", line 496, in _call_chain result = func(*args) File "/Users/desmondlim/.pyenv/versions/3.10.4/lib/python3.10/urllib/request.py", line 643, in http_error_default raise HTTPError(req.full_url, code, msg, hdrs, fp) urllib.error.HTTPError: HTTP Error 401: Unauthorized
错误原因
- 未受保护的URL请求:第38行
print(urllib.request.urlopen(url_entry["url"]).status)直接调用urlopen,没有被try-except包裹,遇到HTTP错误时直接抛出异常终止程序。 - 重复请求URL:脚本先请求一次URL打印状态,然后在
try块里又请求一次,既浪费资源又增加不必要的延迟。 - 错误处理不完整:原来的
except块仅打印错误信息,没有对错误状态码进行计数统计,且处理完错误后直接continue,跳过后续逻辑。
修改方案
关键修改点
- 移除无保护的
urlopen调用,所有URL请求都放到try-except块中处理。 - 在
HTTPError捕获块中获取错误状态码,对应更新计数。 - 合并URL请求逻辑,避免重复请求。
- 修复循环结束后统计信息的打印逻辑,确保输出所有URL的统计数据。
修改后的完整代码
import time import urllib.request import urllib.error url_data = [] total_duration = 0.5 wait_time = 5 counter = 1 # 读取CSV文件 with open("url_list.csv", "r") as url_file: headers = next(url_file).strip().replace(" ","").split(",") for row in url_file: url_list = row.strip().replace(" ","").split(",") # 初始化每个URL的计数为0 url_dict = dict(zip(headers, url_list)) url_dict.update({ "http200": 0, "http300": 0, "http400": 0, "http500": 0 }) url_data.append(url_dict) interval = time.time() + 60 * total_duration print(f"初始探测结束时间戳: {interval}") while True: print(f"\n=== 第 {counter} 轮探测 ===") for url_entry in url_data: print(f"网站名称: {url_entry['name']}") print(f"目标URL: {url_entry['url']}") request_status = None try: with urllib.request.urlopen(url_entry["url"]) as response: request_status = response.status print(f"响应状态码: {request_status}") print("请求成功!") except urllib.error.HTTPError as error: request_status = error.status print(f"HTTP错误: {request_status} - {error.reason}") except urllib.error.URLError as error: print(f"URL错误: {error.reason}") # 网络类错误不统计状态码,直接进入下一个URL print(f"{url_entry['name']} 统计: 200={url_entry['http200']} | 300={url_entry['http300']} | 400={url_entry['http400']} | 500={url_entry['http500']}\n") continue # 统计状态码 if request_status: if 200 <= request_status <= 299: url_entry["http200"] += 1 elif 300 <= request_status <= 399: url_entry["http300"] += 1 elif 400 <= request_status <= 499: url_entry["http400"] += 1 elif 500 <= request_status <= 599: url_entry["http500"] += 1 # 打印当前统计 print(f"{url_entry['name']} 统计: 200={url_entry['http200']} | 300={url_entry['http300']} | 400={url_entry['http400']} | 500={url_entry['http500']}\n") time.sleep(wait_time) counter += 1 # 检查是否到达时间间隔,输出汇总 if time.time() >= interval: print(f"\n=== 过去 {60 * total_duration} 秒内共完成 {counter-1} 轮探测 ===") for url_entry in url_data: print(f"{url_entry['name']} 最终统计: 200={url_entry['http200']} | 300={url_entry['http300']} | 400={url_entry['http400']} | 500={url_entry['http500']}") # 更新下一轮结束时间戳 interval = time.time() + 60 * total_duration counter = 1
修改说明
- 移除无保护请求:删掉了原来直接调用
urlopen的代码,所有请求都在try-except中处理,确保异常不会终止程序。 - 错误状态码统计:在捕获
HTTPError时获取错误状态码,同样进行计数,不会遗漏错误请求的统计。 - 避免重复请求:只请求一次URL,同时完成状态获取和错误处理,提升效率。
- 完善统计输出:时间间隔结束后遍历所有URL条目打印统计,而不是只打印最后一个。
内容的提问来源于stack exchange,提问作者Ldd
相关产品推荐
相关产品推荐

