Google Colab中使用requests访问rasm.io等网站失败如何解决?
排查步骤
首先先执行以下代码捕获具体错误信息,定位问题根因:
import requests import urllib3 # 可选:关闭SSL警告 urllib3.disable_warnings() url = "https://rasm.io/" try: resp = requests.get(url, timeout=15) print(f"请求成功,状态码:{resp.status_code}") print(f"响应前1000字符:\n{resp.text[:1000]}") except Exception as e: print(f"错误类型:{type(e).__name__}") print(f"错误详情:{str(e)}")
同时可以在Colab单元格执行系统命令验证原生网络连通性:
!curl -m 10 -v https://rasm.io/
对应解决方案
根据捕获到的错误类型对应处理:
- SSL证书验证错误:请求时添加
verify=False参数跳过证书校验即可resp = requests.get(url, timeout=15, verify=False) - 连接超时/连接被拒绝:属于Google Colab出口IP被目标站点的WAF/防火墙封禁,或是目标站点屏蔽了数据中心类IP段,给Colab配置可用代理后再发起请求即可。
- 403 Forbidden错误:目标站点配置了反爬策略(常见为Cloudflare五秒盾校验),仅改UA无法绕过,可使用
cloudscraper库替代requests发起请求:# 安装依赖 !pip install -q cloudscraper import cloudscraper scraper = cloudscraper.create_scraper() resp = scraper.get(url, timeout=15) print(resp.text) - 重定向次数过多错误:添加
allow_redirects=False参数查看中间跳转流程,确认是否存在循环重定向,再针对性补充跳转需要的校验参数即可。 - 若curl命令也无法获取响应,说明Colab当前节点本身无法访问该站点,更换运行环境即可解决。
内容的提问来源于stack exchange,提问作者hamidreza bina
相关产品推荐
相关产品推荐

