You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python下载含aspx/ashx后缀的网页文件问题求助

解决ashx/aspx后缀链接的文件下载问题

你遇到的问题是:普通的requests下载代码能处理常规文件链接,但面对ashx/aspx这类动态生成的下载链接时无法正常工作,比如示例链接https://kingcounty.gov/~/media/depts/health/communicable-diseases/documents/C19/data/vax_public.ashx?la=en。

原可正常运行的代码:

import requests
 
print('Download Starting...')
 
url = 'http://www.tutorialspoint.com/python3/python_tutorial.pdf'
 
r = requests.get(url)
 
filename = url.split('/')[-1] # this will take only -1 splitted part of the url
 
with open(filename,'wb') as output_file:
    output_file.write(r.content)
 
print('Download Completed!!!')

修改方案及原因

这类动态链接的服务器通常会做以下限制,针对性调整即可解决:

  • 服务器会校验请求是否来自真实浏览器,拒绝无请求头的爬虫请求
  • 动态链接可能会触发重定向,需要确保请求跟进跳转
  • 原代码通过URL分割获取文件名的方式,在这类带参数的链接里会得到vax_public.ashx?la=en这种无效文件名,真实文件名通常藏在响应头的Content-Disposition字段中

修改后的完整代码

import requests
from urllib.parse import unquote

def download_file(url):
    # 模拟浏览器请求头,避免被服务器拦截
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    }

    print('Download Starting...')
    # 发送GET请求,允许自动跟进重定向
    r = requests.get(url, headers=headers, allow_redirects=True)
    # 检查请求是否成功,失败则抛出异常
    r.raise_for_status()

    # 从响应头解析真实文件名
    filename = None
    if 'Content-Disposition' in r.headers:
        disp = r.headers['Content-Disposition']
        # 提取filename参数,处理URL编码的情况
        if 'filename=' in disp:
            filename = disp.split('filename=')[-1].strip('"')
            filename = unquote(filename)  # 解码URL编码的文件名
    # 如果没拿到真实文件名, fallback到URL分割的方式(去除参数部分)
    if not filename:
        filename = url.split('/')[-1].split('?')[0]

    # 写入文件
    with open(filename, 'wb') as output_file:
        output_file.write(r.content)
    
    print(f'Download Completed!!! 文件已保存为: {filename}')

# 测试示例链接
download_url = 'https://kingcounty.gov/~/media/depts/health/communicable-diseases/documents/C19/data/vax_public.ashx?la=en'
download_file(download_url)

关键修改点说明

  • 添加User-Agent请求头,伪装成Chrome浏览器,绕过服务器的爬虫检测
  • 明确启用allow_redirects=True(requests默认开启,此处写出更清晰),确保跟进动态链接的重定向
  • 新增解析Content-Disposition响应头的逻辑,获取服务器返回的真实文件名,避免生成无效的ashx后缀文件
  • 加入r.raise_for_status(),请求失败时直接抛出异常,方便排查问题
  • 使用unquote()解码URL编码的文件名,避免出现乱码

内容的提问来源于stack exchange,提问作者J_Py

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 17:06:49