You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python requests下载Zip文件失败:未知格式且大小异常求助

解决requests下载Zip文件显示未知格式的问题

问题情况

  • 使用Python 3.8.12的requests库下载香港统计处的Zip文件,保存后打开提示「Unknown file format」,所有下载的Zip文件大小固定为18KB,无法正常打开。
  • 更换为直接指向CSV文件的URL时,代码可正常下载并打开文件;尝试过requests、urllib、wget等多个库均无法解决Zip文件的问题。

原因分析

你的推测完全正确:你使用的Zip文件URL并非直接指向文件资源,而是指向一个跳转网页。服务器返回的18KB内容实际是HTML页面,而非真正的Zip文件。这类网站通常会通过网页跳转提供下载链接,而非直接返回文件,部分还会校验请求头来限制非浏览器请求。

解决方案

1. 模拟浏览器请求头

很多网站会校验User-Agent字段判断请求来源,先添加浏览器请求头尝试获取正确响应:

import requests

file_url = 'https://www.censtatd.gov.hk/en/EIndexbySubject.html?pcode=D5600091&scode=300&file=D5600091B2022MM11B.zip'
# 模拟Chrome浏览器的请求头
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

# 发送请求并检查响应类型
response = requests.get(file_url, headers=headers, allow_redirects=True)
print("响应内容类型:", response.headers.get('Content-Type'))

如果输出仍为text/html,说明需要进一步提取页面内的真实下载链接。

2. 解析网页提取真实Zip链接

使用BeautifulSoup解析返回的HTML页面,找到实际的Zip文件下载链接:

from bs4 import BeautifulSoup

# 解析HTML页面
soup = BeautifulSoup(response.text, 'html.parser')
# 查找页面中所有指向Zip文件的链接
zip_links = soup.find_all('a', href=lambda href: href and href.endswith('.zip'))

if zip_links:
    # 补全完整URL(注意域名前缀)
    real_zip_url = 'https://www.censtatd.gov.hk' + zip_links[0]['href']
    # 下载真实的Zip文件
    zip_response = requests.get(real_zip_url, headers=headers, stream=True)
    # 保存文件
    with open('target.zip', 'wb') as f:
        for chunk in zip_response.iter_content(chunk_size=1024):
            f.write(chunk)
    print("Zip文件下载完成")
else:
    print("未找到有效的Zip文件下载链接")

补充说明

CSV文件的URL可以正常下载,是因为该URL直接指向文件资源,服务器收到请求后直接返回文件内容,无需页面跳转或额外校验。

内容的提问来源于stack exchange,提问作者MadMaple

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 22:40:37