使用requests爬取尼泊尔证券网站遇SSL及连接错误的解决咨询
问题描述
我是一名编程初学者,尝试使用requests库爬取https://www.nepalstock.com.np/网站数据时,首次请求出现ssl.SSLCertVerificationError证书验证失败错误;尝试添加verify=False参数后,又触发requests.exceptions.ConnectionError连接断开错误。我曾尝试用certifi引用CA证书,但仍出现初始错误。请问这两个错误是否相关?该如何解决?
首次请求代码
import requests web = requests.get("https://www.nepalstock.com.np/")
对应错误
ssl.SSLCertVerificationError: [SSL: CERTIFICATE_VERIFY_FAILED] certificate verify failed: unable to get local issuer certificate (_ssl.c:997) ... requests.exceptions.SSLError: HTTPSConnectionPool(host='www.nepalstock.com.np', port=443): Max retries exceeded with url: / (Caused by SSLError(...))
添加verify=False后的代码
import requests web = requests.get("https://www.nepalstock.com.np/", verify=False)
对应错误
http.client.RemoteDisconnected: Remote end closed connection without response ... requests.exceptions.ConnectionError: ('Connection aborted.', RemoteDisconnected(...))
解答
这两个错误直接相关,核心问题出在SSL握手环节:
- 初始的证书验证失败,是因为本地CA证书库无法识别目标网站的SSL证书(可能是证书链不完整、本地证书库过时);
- 你用
verify=False跳过验证后,网站服务器检测到不安全的请求(非标准SSL握手),主动断开了连接,所以触发连接错误。
下面是具体的解决步骤:
1. 用最新CA证书验证
先确保certifi是最新版本,它提供了当前最全的CA证书集合:
pip install --upgrade certifi
然后在请求中指定certifi的证书路径:
import requests import certifi response = requests.get("https://www.nepalstock.com.np/", verify=certifi.where())
如果还是不行,手动导出目标网站的根证书:
- 打开浏览器访问该网站,点击地址栏的锁图标 → 证书 → 详细信息 → 复制到文件,选择PEM格式保存;
- 在代码中指定这个证书文件的路径:
import requests response = requests.get("https://www.nepalstock.com.np/", verify="/path/to/your/exported-cert.pem")
2. 伪装浏览器请求头
目标网站可能会拦截非浏览器发起的请求,即使SSL验证通过也会被拒。添加浏览器的User-Agent头:
import requests import certifi headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } response = requests.get("https://www.nepalstock.com.np/", headers=headers, verify=certifi.where())
3. 排查网络环境
如果你在代理、VPN或公司防火墙下,这些设备可能会篡改SSL证书导致验证失败。尝试切换到普通网络,或者配置requests使用正确的代理:
import requests import certifi proxies = { "http": "http://your-proxy-address:port", "https": "https://your-proxy-address:port" } response = requests.get("https://www.nepalstock.com.np/", proxies=proxies, verify=certifi.where())
重要提醒
- 绝对不要在生产环境中使用
verify=False,这会让你的请求完全暴露在中间人攻击的风险中; - 该网站可能有反爬机制,频繁请求会被封禁IP,建议每次请求后添加1-2秒的间隔。
内容的提问来源于stack exchange,提问作者pawan kharel
相关产品推荐
相关产品推荐

