You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让Python脚本遇urllib.request HTTP错误时继续执行?

解决URL探测脚本遇HTTP错误终止的问题

我有个读取URL列表CSV的Python脚本,CSV内容如下:

name, url
google, https://httpstat.us/200
yahoo, https://httpstat.us/401
bcs, https://httpstat.us/521

当前脚本运行时,遇到返回HTTP错误码的URL就直接终止,无法处理剩余URL,报错信息如下:

➜  DevOps_Practice python urlProbe.py
1687386788.8879528
Counter: 1
Website: google
URL: https://httpstat.us/200
200
200
All is good!
 google - 1 - 0 - 0 - 0

Website: yahoo
URL: https://httpstat.us/401
Traceback (most recent call last):
  File "/Users/desmondlim/Documents/DevOps/Projects/DevOps_Challenge_SPH/urlProbe.py", line 38, in <module>
    print(urllib.request.urlopen(url_entry["url"]).status)
  File "/Users/desmondlim/.pyenv/versions/3.10.4/lib/python3.10/urllib/request.py", line 216, in urlopen
    return opener.open(url, data, timeout)
  File "/Users/desmondlim/.pyenv/versions/3.10.4/lib/python3.10/urllib/request.py", line 525, in open
    response = meth(req, response)
  File "/Users/desmondlim/.pyenv/versions/3.10.4/lib/python3.10/urllib/request.py", line 634, in http_response
    response = self.parent.error(
  File "/Users/desmondlim/.pyenv/versions/3.10.4/lib/python3.10/urllib/request.py", line 563, in error
    return self._call_chain(*args)
  File "/Users/desmondlim/.pyenv/versions/3.10.4/lib/python3.10/urllib/request.py", line 496, in _call_chain
    result = func(*args)
  File "/Users/desmondlim/.pyenv/versions/3.10.4/lib/python3.10/urllib/request.py", line 643, in http_error_default
    raise HTTPError(req.full_url, code, msg, hdrs, fp)
urllib.error.HTTPError: HTTP Error 401: Unauthorized

错误原因

  • 未受保护的URL请求:第38行print(urllib.request.urlopen(url_entry["url"]).status)直接调用urlopen,没有被try-except包裹,遇到HTTP错误时直接抛出异常终止程序。
  • 重复请求URL:脚本先请求一次URL打印状态,然后在try块里又请求一次,既浪费资源又增加不必要的延迟。
  • 错误处理不完整:原来的except块仅打印错误信息,没有对错误状态码进行计数统计,且处理完错误后直接continue,跳过后续逻辑。

修改方案

关键修改点

  1. 移除无保护的urlopen调用,所有URL请求都放到try-except块中处理。
  2. 在HTTPError捕获块中获取错误状态码,对应更新计数。
  3. 合并URL请求逻辑,避免重复请求。
  4. 修复循环结束后统计信息的打印逻辑,确保输出所有URL的统计数据。

修改后的完整代码

import time
import urllib.request
import urllib.error

url_data = []
total_duration = 0.5
wait_time = 5
counter = 1

# 读取CSV文件
with open("url_list.csv", "r") as url_file:
    headers = next(url_file).strip().replace(" ","").split(",")
    for row in url_file:
        url_list = row.strip().replace(" ","").split(",")
        # 初始化每个URL的计数为0
        url_dict = dict(zip(headers, url_list))
        url_dict.update({
            "http200": 0,
            "http300": 0,
            "http400": 0,
            "http500": 0
        })
        url_data.append(url_dict)

interval = time.time() + 60 * total_duration
print(f"初始探测结束时间戳: {interval}")

while True:
    print(f"\n=== 第 {counter} 轮探测 ===")
    for url_entry in url_data:
        print(f"网站名称: {url_entry['name']}")
        print(f"目标URL: {url_entry['url']}")
        request_status = None
        
        try:
            with urllib.request.urlopen(url_entry["url"]) as response:
                request_status = response.status
                print(f"响应状态码: {request_status}")
                print("请求成功!")
        except urllib.error.HTTPError as error:
            request_status = error.status
            print(f"HTTP错误: {request_status} - {error.reason}")
        except urllib.error.URLError as error:
            print(f"URL错误: {error.reason}")
            # 网络类错误不统计状态码,直接进入下一个URL
            print(f"{url_entry['name']} 统计: 200={url_entry['http200']} | 300={url_entry['http300']} | 400={url_entry['http400']} | 500={url_entry['http500']}\n")
            continue
        
        # 统计状态码
        if request_status:
            if 200 <= request_status <= 299:
                url_entry["http200"] += 1
            elif 300 <= request_status <= 399:
                url_entry["http300"] += 1
            elif 400 <= request_status <= 499:
                url_entry["http400"] += 1
            elif 500 <= request_status <= 599:
                url_entry["http500"] += 1
        
        # 打印当前统计
        print(f"{url_entry['name']} 统计: 200={url_entry['http200']} | 300={url_entry['http300']} | 400={url_entry['http400']} | 500={url_entry['http500']}\n")
    
    time.sleep(wait_time)
    counter += 1
    
    # 检查是否到达时间间隔,输出汇总
    if time.time() >= interval:
        print(f"\n=== 过去 {60 * total_duration} 秒内共完成 {counter-1} 轮探测 ===")
        for url_entry in url_data:
            print(f"{url_entry['name']} 最终统计: 200={url_entry['http200']} | 300={url_entry['http300']} | 400={url_entry['http400']} | 500={url_entry['http500']}")
        # 更新下一轮结束时间戳
        interval = time.time() + 60 * total_duration
        counter = 1

修改说明

  1. 移除无保护请求:删掉了原来直接调用urlopen的代码,所有请求都在try-except中处理,确保异常不会终止程序。
  2. 错误状态码统计:在捕获HTTPError时获取错误状态码,同样进行计数,不会遗漏错误请求的统计。
  3. 避免重复请求:只请求一次URL,同时完成状态获取和错误处理,提升效率。
  4. 完善统计输出:时间间隔结束后遍历所有URL条目打印统计,而不是只打印最后一个。

内容的提问来源于stack exchange,提问作者Ldd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 05:55:00