You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python与BeautifulSoup:下载/download结尾链接文件如何获取原始文件名?

解决方法

当访问以/download结尾的链接时,服务器通常会在响应头的Content-Disposition字段中返回文件的原始名称——这也是手动下载时浏览器能获取正确文件名的核心原因。我们可以通过解析这个字段拿到文件名和扩展名,再完成下载。

方法一:使用requests库(推荐,代码更简洁)

先确保安装依赖:pip install requests

修改后的代码如下:

from bs4 import BeautifulSoup
import requests
import os
import re

# 提前创建保存目录,避免报错
os.makedirs('pics', exist_ok=True)

with open("page.html") as html_page:
    soup = BeautifulSoup(html_page, 'lxml')

# 直接筛选出所有下载链接,避免后续二次判断
download_links = [
    link.get('href') 
    for link in soup.findAll('a') 
    if link.get('href') and link.get('href').endswith('download')
]

for idx, url in enumerate(download_links):
    try:
        # 发起流式请求,避免大文件占用过多内存
        response = requests.get(url, stream=True)
        response.raise_for_status()  # 捕获请求错误

        # 解析Content-Disposition字段提取文件名
        filename = None
        if 'Content-Disposition' in response.headers:
            content_disposition = response.headers['Content-Disposition']
            # 兼容带引号和不带引号的文件名格式
            match = re.search(r'filename[^;=\n]*=(([""]).*?\2|[^;\n]*)', content_disposition)
            if match:
                filename = match.group(1).strip('"')
        
        # 若无法获取原始文件名,用索引作为兜底命名
        if not filename:
            filename = f'file_{idx}.unknown'
        
        # 保存文件
        save_path = os.path.join('pics', filename)
        with open(save_path, 'wb') as f:
            for chunk in response.iter_content(chunk_size=8192):
                f.write(chunk)
        print(f"已保存: {save_path}")
    except Exception as e:
        print(f"下载失败 {url}: {str(e)}")

方法二:使用原生urllib库(无需额外安装)

如果不想引入第三方库,可直接用Python标准库实现:

from bs4 import BeautifulSoup
import urllib.request
import os
import re

os.makedirs('pics', exist_ok=True)

with open("page.html") as html_page:
    soup = BeautifulSoup(html_page, 'lxml')

download_links = [
    link.get('href') 
    for link in soup.findAll('a') 
    if link.get('href') and link.get('href').endswith('download')
]

for idx, url in enumerate(download_links):
    try:
        with urllib.request.urlopen(url) as response:
            headers = response.info()
            filename = None

            # 解析响应头中的文件名
            if 'Content-Disposition' in headers:
                content_disposition = headers['Content-Disposition']
                match = re.search(r'filename[^;=\n]*=(([""]).*?\2|[^;\n]*)', content_disposition)
                if match:
                    filename = match.group(1).strip('"')
            
            if not filename:
                filename = f'file_{idx}.unknown'
            
            save_path = os.path.join('pics', filename)
            with open(save_path, 'wb') as f:
                f.write(response.read())
            print(f"已保存: {save_path}")
    except Exception as e:
        print(f"下载失败 {url}: {str(e)}")

关键细节说明

  • 正则表达式兼容两种常见的Content-Disposition格式:attachment; filename="example.jpg"和attachment; filename=example.pdf。
  • 用enumerate替代links.index(item),避免重复链接导致的索引错误,同时提升效率。
  • 加入异常处理,避免单个链接下载失败导致整个脚本中断。
  • 提前创建保存目录,防止文件保存时因目录不存在报错。

内容的提问来源于stack exchange,提问作者murd0x

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 04:50:25