使用urllib或requests下载文件遇403报错无法获取csv文件求助
问题原因
- 你使用的URL是网页展示地址,不是CSV文件的真实下载路径,所以下载的内容是HTML网页源码,不是你要的CSV数据
- 403报错是因为目标站点有基础反爬校验,未携带
User-Agent标识的爬虫请求会被直接拦截
修复后代码
import requests # 替换为你从网页复制的CSV真实下载链接 remote_url = 'https://hdr.undp.org/sites/default/files/2023_statistical_update_137506.csv' # 添加请求头伪装浏览器请求,解决403问题 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } local_file = 'human_development_index_data.csv' response = requests.get(remote_url, headers=headers) # 确认请求成功再写入文件 response.raise_for_status() with open(local_file, 'wb') as file: file.write(response.content)
注意事项
如果代码中的直链失效,可手动打开目标网页,找到CSV下载按钮后右键选择「复制链接地址」,替换代码中remote_url的取值即可,保留请求头配置即可避免403报错。
内容的提问来源于stack exchange,提问作者RThatha
相关产品推荐
相关产品推荐

