使用Python requests下载无后缀动态URL的Excel文件失败如何解决
Python requests下载无后缀动态Excel文件失败的解决方案
问题原因
- 目标站点配置了基础反爬校验,未携带正常浏览器请求头的访问会被拦截,返回错误提示页面的HTML内容
- 部分站点要求请求携带访问前置页面生成的Cookie,直接请求下载地址会被判定为非法访问
修复方案
方案1:添加基础请求头(适合无Cookie校验的站点)
import requests # 模拟Chrome浏览器的请求头,可替换为自己浏览器的实际UA headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } download_url = 'https://www.djppr.kemenkeu.go.id/page/loadViewer?idViewer=9369&action=download' response = requests.get(download_url, headers=headers, timeout=30, allow_redirects=True) # 校验请求是否成功 response.raise_for_status() with open('file.xlsx', 'wb') as f: f.write(response.content)
方案2:使用会话留存Cookie(适合有前置页面Cookie校验的站点)
import requests headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Referer": "https://www.djppr.kemenkeu.go.id/" # 配置站点域名作为来源页,降低被拦截概率 } # 初始化会话自动管理Cookie session = requests.Session() session.headers.update(headers) # 先访问站点主页获取合法Cookie session.get("https://www.djppr.kemenkeu.go.id/", timeout=30) # 再请求下载地址 download_resp = session.get('https://www.djppr.kemenkeu.go.id/page/loadViewer?idViewer=9369&action=download', timeout=30) download_resp.raise_for_status() with open('file.xlsx', 'wb') as f: f.write(download_resp.content)
排查技巧
- 如果修改后仍返回HTML内容,可打印
download_resp.text查看返回的页面内容,判断是验证码拦截、权限不足还是其他校验要求,对应调整请求参数即可 - 可以通过
download_resp.headers.get('Content-Type')校验返回内容类型:Excel文件的Content-Type通常为application/vnd.openxmlformats-officedocument.spreadsheetml.sheet或application/vnd.ms-excel,若为text/html则说明仍被拦截 - 无需强制指定文件名后缀,可以从响应头的
Content-Disposition字段提取官方给定的文件名,避免后缀错误
内容的提问来源于stack exchange,提问作者Vidy45agarvk
相关产品推荐
相关产品推荐

