You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Requests模块无法下载目标Excel文件的问题求助

问题描述

手动访问EIA的Excel文件链接可正常下载,但使用以下Requests代码无法获取有效文件:

import requests
     
url = 'https://www.eia.gov/electricity/data/eia860m/archive/xls/september_generator2019.xlsx'
response = requests.get(url)
response.raise_for_status()
with open(path, 'wb') as outfile:
    outfile.write(response.content)

生成的文件无法被Excel打开,查看response.history发现请求被重定向至EIA的另一页面。使用pandas直接读取也报错,需求是仅用Requests模块解决该问题。

原因分析
  • 目标网站存在基础反爬机制,会校验请求的User-Agent头部:Requests库默认的请求头没有模拟真实浏览器的标识,服务器识别出非浏览器请求后,会将请求重定向到其他页面,而非返回真实的Excel文件。
  • 代码中保存的response.content实际是重定向后的页面HTML内容,并非Excel二进制数据,因此无法被Excel识别。
修复方案

在请求中添加模拟浏览器的User-Agent头部,让服务器认为请求来自真实浏览器,即可获取到正确的Excel文件。修改后的代码如下:

import requests

url = 'https://www.eia.gov/electricity/data/eia860m/archive/xls/september_generator2019.xlsx'
# 添加模拟浏览器的请求头
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

response = requests.get(url, headers=headers)
response.raise_for_status()
# 可选:验证响应是否为Excel文件
if 'application/vnd.openxmlformats-officedocument.spreadsheetml.sheet' in response.headers.get('Content-Type', ''):
    with open('september_generator2019.xlsx', 'wb') as outfile:
        outfile.write(response.content)
    print("Excel文件下载成功")
else:
    print("未获取到有效Excel文件,响应内容类型:", response.headers.get('Content-Type'))

关键修改说明

  • 添加headers参数,传入包含User-Agent的字典,模拟Chrome浏览器的请求标识,绕过基础反爬校验。
  • 可选添加Content-Type校验,确保下载的内容确实是Excel文件,避免再次保存错误内容。

内容的提问来源于stack exchange,提问作者user6794223

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 08:05:25