使用Requests库无法从链接下载xls文件:陷入循环或超时
问题描述
尝试从以下链接下载文件:
url = "https://nsearchives.nseindia.com/content/fo/qtyfreeze.xls"
已尝试设置timeout、stream、headers、allow-redirects参数,但遇到以下问题:
- 未设置
timeout时,响应始终无法完成 - 设置
timeout后,一直触发超时错误
但在Firefox和Chrome浏览器中可正常下载该文件。
使用的代码如下:
import requests as req headers = { 'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36', "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8,application/vnd.ms-excel", "Accept-Encoding": "gzip, deflate"} url = "https://nsearchives.nseindia.com/content/fo/qtyfreeze.xls" # 以下代码尝试过多种参数组合 resp = req.get(url, timeout=30, headers=headers) # 也试过urllib #urllib.request.urlretrieve(url,"/home/data/freeze_quantity.xls")
解决方案
NSE的档案服务器会对非浏览器请求做额外校验,仅添加常规请求头不足以通过验证,需要补充关键请求头并优化下载方式:
- 补充
Referer和Host头,模拟浏览器的完整请求上下文 - 开启
stream=True分块下载,避免因单次读取大文件导致超时 - 适当延长超时时间(比如60秒)
修改后的Requests代码示例
import requests as req headers = { 'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36', "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8,application/vnd.ms-excel", "Accept-Encoding": "gzip, deflate", "Referer": "https://www.nseindia.com/", "Host": "nsearchives.nseindia.com" } url = "https://nsearchives.nseindia.com/content/fo/qtyfreeze.xls" try: # 开启流模式并设置较长超时 with req.get(url, timeout=60, headers=headers, stream=True) as resp: resp.raise_for_status() # 检查请求是否成功 with open("qtyfreeze.xls", "wb") as f: # 分块写入文件,每块1024字节 for chunk in resp.iter_content(chunk_size=1024): f.write(chunk) print("文件下载完成") except req.exceptions.RequestException as e: print(f"下载出错: {e}")
修改后的Urllib代码示例
import urllib.request headers = { 'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36', "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8,application/vnd.ms-excel", "Accept-Encoding": "gzip, deflate", "Referer": "https://www.nseindia.com/", "Host": "nsearchives.nseindia.com" } req = urllib.request.Request(url, headers=headers) with urllib.request.urlopen(req, timeout=60) as response, open("/home/data/freeze_quantity.xls", "wb") as out_file: out_file.write(response.read())
核心原因是NSE的档案服务器会验证请求的Referer(来源页面)和Host(目标主机),缺少这些头会被服务器限流或静默拒绝响应,最终导致超时。
内容的提问来源于stack exchange,提问作者Sid
相关产品推荐
相关产品推荐

