使用Requests爬取URL时为何触发SSLError - [SSL: UNEXPECTED_EOF_WHILE_READING]?
问题解决:爬取网页时SSLError未被捕获导致程序中断
问题背景
安装指定依赖并编写爬取代码后,尝试捕获连接异常但仍遇到SSL错误导致程序中断:
安装依赖命令
!pip install httpx !pip install selectolax
import httpx from selectolax.parser import HTMLParser
原始爬取代码
import requests urls = ['http://toofab.com/2017/05/08/real-housewives-atlanta-kandi-burruss-rape-phaedra-parks-porsha-williams/', 'https://www.today.com/style/see-people-s-choice-awards-red-carpet-looks-t141832', 'https://www.zerchoo.com/entertainment/gossip-girl-10-years-later-how-upper-east-siders-shocked-the-world-changed-pop-culture-forever/', 'www.intouchweekly.com/posts/gwen-stefani-dumped-156076'] text_column = [] for url_index in range(len(urls)): try: resp= requests.get(urls[url_index], headers={ 'User-Agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36' } ) html = HTMLParser(resp.text) article = [element.text().strip() for element in html.css('p')] article_text = ''.join(map(str,article)) text_column.append(article_text) except (httpx.ConnectError, httpx.HTTPStatusError): print(f"Connection error for URL: {urls[url_index]}") continue
触发的错误
SSLError: HTTPSConnectionPool(host='www.zerchoo.com', port=443): Max retries exceeded with url: /entertainment/gossip-girl-10-years-later-how-upper-east-siders-shocked-the-world-changed-pop-culture-forever/ (Caused by SSLError(SSLEOFError(8, '[SSL: UNEXPECTED_EOF_WHILE_READING] EOF occurred in violation of protocol (_ssl.c:1007)')))
问题原因
你捕获的是httpx库的异常,但实际发送请求用的是requests库,抛出的SSLError属于requests.exceptions.SSLError,不在当前捕获的异常类型范围内,所以程序直接中断。另外第四个URL缺少http://或https://前缀,也会触发格式错误异常。
解决方案
- 调整异常捕获范围,加入
requests相关的异常类型,比如用requests.exceptions.RequestException涵盖所有请求类异常; - 处理URL前缀问题,确保每个URL都有合法的协议头;
- 优化循环写法,直接遍历URL更简洁直观。
修改后的完整代码
import requests from selectolax.parser import HTMLParser def add_url_prefix(url): # 为缺少协议头的URL补全http前缀 if not url.startswith(('http://', 'https://')): return f'http://{url}' return url urls = ['http://toofab.com/2017/05/08/real-housewives-atlanta-kandi-burruss-rape-phaedra-parks-porsha-williams/', 'https://www.today.com/style/see-people-s-choice-awards-red-carpet-looks-t141832', 'https://www.zerchoo.com/entertainment/gossip-girl-10-years-later-how-upper-east-siders-shocked-the-world-changed-pop-culture-forever/', 'www.intouchweekly.com/posts/gwen-stefani-dumped-156076'] text_column = [] for url in urls: processed_url = add_url_prefix(url) try: resp = requests.get(processed_url, headers={ 'User-Agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36' } ) # 主动触发HTTP状态码异常(4xx/5xx) resp.raise_for_status() html = HTMLParser(resp.text) article = [element.text().strip() for element in html.css('p')] article_text = ''.join(article) text_column.append(article_text) except requests.exceptions.RequestException as e: print(f"处理URL {processed_url} 时出错: {str(e)}") continue
补充说明
requests.exceptions.RequestException会捕获requests库所有请求相关异常,包括SSL错误、连接超时、HTTP状态码错误等;add_url_prefix函数解决了无协议头URL的格式问题;resp.raise_for_status()会在请求返回错误状态码时主动抛出异常,统一进入异常处理分支。
内容的提问来源于stack exchange,提问作者Hedra Lotfy
相关产品推荐
相关产品推荐

