You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Requests爬取URL时为何触发SSLError - [SSL: UNEXPECTED_EOF_WHILE_READING]?

问题解决:爬取网页时SSLError未被捕获导致程序中断

问题背景

安装指定依赖并编写爬取代码后,尝试捕获连接异常但仍遇到SSL错误导致程序中断:

安装依赖命令

!pip install httpx
!pip install selectolax
import httpx
from selectolax.parser import HTMLParser

原始爬取代码

import requests

urls = ['http://toofab.com/2017/05/08/real-housewives-atlanta-kandi-burruss-rape-phaedra-parks-porsha-williams/', 
        'https://www.today.com/style/see-people-s-choice-awards-red-carpet-looks-t141832', 
        'https://www.zerchoo.com/entertainment/gossip-girl-10-years-later-how-upper-east-siders-shocked-the-world-changed-pop-culture-forever/', 
        'www.intouchweekly.com/posts/gwen-stefani-dumped-156076']
text_column = []
for url_index in range(len(urls)):
  try:
    resp= requests.get(urls[url_index],
        headers={
        'User-Agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36'
        }
    )
    html = HTMLParser(resp.text)
    article = [element.text().strip() for element in html.css('p')]
    article_text = ''.join(map(str,article))
    text_column.append(article_text)
  except (httpx.ConnectError, httpx.HTTPStatusError):
    print(f"Connection error for URL: {urls[url_index]}")
    continue

触发的错误

SSLError: HTTPSConnectionPool(host='www.zerchoo.com', port=443): Max retries exceeded with url: /entertainment/gossip-girl-10-years-later-how-upper-east-siders-shocked-the-world-changed-pop-culture-forever/ (Caused by SSLError(SSLEOFError(8, '[SSL: UNEXPECTED_EOF_WHILE_READING] EOF occurred in violation of protocol (_ssl.c:1007)')))

问题原因

你捕获的是httpx库的异常,但实际发送请求用的是requests库,抛出的SSLError属于requests.exceptions.SSLError,不在当前捕获的异常类型范围内,所以程序直接中断。另外第四个URL缺少http://或https://前缀,也会触发格式错误异常。

解决方案

  1. 调整异常捕获范围,加入requests相关的异常类型,比如用requests.exceptions.RequestException涵盖所有请求类异常;
  2. 处理URL前缀问题,确保每个URL都有合法的协议头;
  3. 优化循环写法,直接遍历URL更简洁直观。

修改后的完整代码

import requests
from selectolax.parser import HTMLParser

def add_url_prefix(url):
    # 为缺少协议头的URL补全http前缀
    if not url.startswith(('http://', 'https://')):
        return f'http://{url}'
    return url

urls = ['http://toofab.com/2017/05/08/real-housewives-atlanta-kandi-burruss-rape-phaedra-parks-porsha-williams/', 
        'https://www.today.com/style/see-people-s-choice-awards-red-carpet-looks-t141832', 
        'https://www.zerchoo.com/entertainment/gossip-girl-10-years-later-how-upper-east-siders-shocked-the-world-changed-pop-culture-forever/', 
        'www.intouchweekly.com/posts/gwen-stefani-dumped-156076']
text_column = []

for url in urls:
    processed_url = add_url_prefix(url)
    try:
        resp = requests.get(processed_url,
            headers={
                'User-Agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36'
            }
        )
        # 主动触发HTTP状态码异常(4xx/5xx)
        resp.raise_for_status()
        html = HTMLParser(resp.text)
        article = [element.text().strip() for element in html.css('p')]
        article_text = ''.join(article)
        text_column.append(article_text)
    except requests.exceptions.RequestException as e:
        print(f"处理URL {processed_url} 时出错: {str(e)}")
        continue

补充说明

  • requests.exceptions.RequestException会捕获requests库所有请求相关异常,包括SSL错误、连接超时、HTTP状态码错误等;
  • add_url_prefix函数解决了无协议头URL的格式问题;
  • resp.raise_for_status()会在请求返回错误状态码时主动抛出异常,统一进入异常处理分支。

内容的提问来源于stack exchange,提问作者Hedra Lotfy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 22:37:02