You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取aaii站点XLS文件出现403 Forbidden错误如何解决

Python读取AAII网站XLS文件报403错误的解决方法

目标网站使用Incapsula安全防护服务,会拦截无有效浏览器标识的爬虫请求,你之前尝试的三种方法都没有携带合法的请求头,因此被判定为非法访问触发403。

可通过添加模拟浏览器的请求头解决该问题,以下是两种可直接运行的方案:


方案1:无需保存本地,直接读取内存中的文件内容

适合不需要留存源文件的场景,减少IO操作:

import pandas as pd
import requests
from io import BytesIO

# 目标文件地址和工作表名
url = 'https://www.aaii.com/files/surveys/sentiment.xls'
sheet_name = 'SENTIMENT'

# 配置请求头,模拟Chrome浏览器访问
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8'
}

# 发送带请求头的GET请求
response = requests.get(url, headers=headers, allow_redirects=True)
# 校验请求是否成功,有异常会主动抛出
response.raise_for_status()

# 直接从内存读取Excel内容
df = pd.read_excel(BytesIO(response.content), sheet_name=sheet_name)
print(df.head())

方案2:先保存到本地再读取

适合需要留存源文件的场景:

import pandas as pd
import requests

url = 'https://www.aaii.com/files/surveys/sentiment.xls'
sheet_name = 'SENTIMENT'
save_path = './sentiment.xls'

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8'
}

response = requests.get(url, headers=headers, allow_redirects=True)
response.raise_for_status()

# 保存到本地
with open(save_path, 'wb') as f:
    f.write(response.content)

# 读取本地文件
df = pd.read_excel(save_path, sheet_name=sheet_name)
print(df.head())

注意事项

如果后续仍然触发拦截,可以将User-Agent替换为你当前浏览器的实际标识,不要短时间内频繁发送请求避免被IP限流。


内容的提问来源于stack exchange,提问作者itsergiu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 16:36:03