访问纳斯达克RSS Feed时Feedparser出现SSL读取操作错误
问题描述
使用Python 3.12 + Feedparser 6.0.11,已安装ca-certificates,访问纳斯达克RSS Feed(https://www.nasdaq.com/feed/rssoutbound?category=Financial+Advisors)时触发KeyboardInterrupt,错误栈如下:
https://www.nasdaq.com/feed/rssoutbound?category=Innovation ^CTraceback (most recent call last): File "/home/nckr/kiwichi/read.py", line 78, in <module> NewsFeed = feedparser.parse(url) ^^^^^^^^^^^^^^^^^^^^^ File "/home/nckr/kiwichi/venv/lib/python3.12/site-packages/feedparser/api.py", line 216, in parse data = _open_resource(url_file_stream_or_string, etag, modified, agent, referrer, handlers, request_headers, result) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/nckr/kiwichi/venv/lib/python3.12/site-packages/feedparser/api.py", line 115, in _open_resource return http.get(url_file_stream_or_string, etag, modified, agent, referrer, handlers, request_headers, result) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/nckr/kiwichi/venv/lib/python3.12/site-packages/feedparser/http.py", line 171, in get f = opener.open(request) ^^^^^^^^^^^^^^^^^^^^ File "/usr/lib/python3.12/urllib/request.py", line 515, in open response = self._open(req, data) ^^^^^^^^^^^^^^^^^^^^^ File "/usr/lib/python3.12/urllib/request.py", line 532, in _open result = self._call_chain(self.handle_open, protocol, protocol + ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/lib/python3.12/urllib/request.py", line 492, in _call_chain result = func(*args) ^^^^^^^^^^^ File "/usr/lib/python3.12/urllib/request.py", line 1392, in https_open return self.do_open(http.client.HTTPSConnection, req, ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/lib/python3.12/urllib/request.py", line 1348, in do_open r = h.getresponse() ^^^^^^^^^^^^^^^ File "/usr/lib/python3.12/http/client.py", line 1428, in getresponse response.begin() File "/usr/lib/python3.12/http/client.py", line 331, in begin version, status, reason = self._read_status() ^^^^^^^^^^^^^^^^^^^ File "/usr/lib/python3.12/http/client.py", line 292, in _read_status line = str(self.fp.readline(_MAXLINE + 1), "iso-8859-1") ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/lib/python3.12/socket.py", line 707, in readinto return self._sock.recv_into(b) ^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/lib/python3.12/ssl.py", line 1252, in recv_into return self.read(nbytes, buffer) ^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/lib/python3.12/ssl.py", line 1104, in read return self._sslobj.read(len, buffer) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ KeyboardInterrupt
已尝试全局设置未验证SSL上下文,但无效:
if hasattr(ssl, '_create_unverified_context'): ssl._create_default_https_context = ssl._create_unverified_context
curl可正常访问该URL,需解决以下问题:
- 如何修复Feedparser的访问超时/SSL问题?
- 如何获取更详细的错误信息,而非仅KeyboardInterrupt?
解决方法
1. 直接设置Feedparser超时(最简洁)
Feedparser 6.0+ 支持在parse方法中直接传入timeout参数,避免因无超时导致手动中断:
import feedparser import ssl # 全局禁用SSL验证(若需要) ssl._create_default_https_context = ssl._create_unverified_context url = "https://www.nasdaq.com/feed/rssoutbound?category=Financial+Advisors" # 设置10秒超时,可根据网络情况调整 feed = feedparser.parse(url, timeout=10) # 检查解析状态 if feed.bozo == 0: print(f"Feed标题: {feed.feed.get('title')}") else: print(f"解析错误详情: {feed.bozo_exception}")
2. 模拟浏览器请求头
纳斯达克可能拦截非浏览器请求,添加User-Agent等头信息伪装请求:
import feedparser import ssl from urllib.request import Request, build_opener, HTTPSHandler url = "https://www.nasdaq.com/feed/rssoutbound?category=Financial+Advisors" headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } # 构建带SSL上下文和请求头的opener ctx = ssl._create_unverified_context() opener = build_opener(HTTPSHandler(context=ctx)) opener.addheaders = list(headers.items()) # 发送带超时的请求 response = opener.open(Request(url), timeout=10) feed = feedparser.parse(response) print(f"Feed条目数: {len(feed.entries)}")
3. 捕获详细异常信息
通过try-except捕获不同类型的异常,精准定位问题:
import feedparser import ssl import socket ssl._create_default_https_context = ssl._create_unverified_context url = "https://www.nasdaq.com/feed/rssoutbound?category=Financial+Advisors" try: feed = feedparser.parse(url, timeout=10) if feed.bozo != 0: print(f"解析异常: {type(feed.bozo_exception).__name__}: {feed.bozo_exception}") else: print("Feed访问解析成功") except socket.timeout: print("请求超时,请检查网络或延长超时时间") except ssl.SSLError as e: print(f"SSL错误: {e}") except Exception as e: print(f"其他错误: {type(e).__name__}: {e}")
关键说明
- 全局SSL设置无效的原因:Feedparser初始化时可能已创建默认handler,后续修改全局上下文不会覆盖已有handler,显式传入上下文更可靠
- curl能访问但Python不行,核心差异是请求头和超时设置,模拟浏览器头+超时基本能解决问题
内容的提问来源于stack exchange,提问作者Dan
相关产品推荐
相关产品推荐

