Python与BeautifulSoup:下载/download结尾链接文件如何获取原始文件名?
解决方法
当访问以/download结尾的链接时,服务器通常会在响应头的Content-Disposition字段中返回文件的原始名称——这也是手动下载时浏览器能获取正确文件名的核心原因。我们可以通过解析这个字段拿到文件名和扩展名,再完成下载。
方法一:使用requests库(推荐,代码更简洁)
先确保安装依赖:pip install requests
修改后的代码如下:
from bs4 import BeautifulSoup import requests import os import re # 提前创建保存目录,避免报错 os.makedirs('pics', exist_ok=True) with open("page.html") as html_page: soup = BeautifulSoup(html_page, 'lxml') # 直接筛选出所有下载链接,避免后续二次判断 download_links = [ link.get('href') for link in soup.findAll('a') if link.get('href') and link.get('href').endswith('download') ] for idx, url in enumerate(download_links): try: # 发起流式请求,避免大文件占用过多内存 response = requests.get(url, stream=True) response.raise_for_status() # 捕获请求错误 # 解析Content-Disposition字段提取文件名 filename = None if 'Content-Disposition' in response.headers: content_disposition = response.headers['Content-Disposition'] # 兼容带引号和不带引号的文件名格式 match = re.search(r'filename[^;=\n]*=(([""]).*?\2|[^;\n]*)', content_disposition) if match: filename = match.group(1).strip('"') # 若无法获取原始文件名,用索引作为兜底命名 if not filename: filename = f'file_{idx}.unknown' # 保存文件 save_path = os.path.join('pics', filename) with open(save_path, 'wb') as f: for chunk in response.iter_content(chunk_size=8192): f.write(chunk) print(f"已保存: {save_path}") except Exception as e: print(f"下载失败 {url}: {str(e)}")
方法二:使用原生urllib库(无需额外安装)
如果不想引入第三方库,可直接用Python标准库实现:
from bs4 import BeautifulSoup import urllib.request import os import re os.makedirs('pics', exist_ok=True) with open("page.html") as html_page: soup = BeautifulSoup(html_page, 'lxml') download_links = [ link.get('href') for link in soup.findAll('a') if link.get('href') and link.get('href').endswith('download') ] for idx, url in enumerate(download_links): try: with urllib.request.urlopen(url) as response: headers = response.info() filename = None # 解析响应头中的文件名 if 'Content-Disposition' in headers: content_disposition = headers['Content-Disposition'] match = re.search(r'filename[^;=\n]*=(([""]).*?\2|[^;\n]*)', content_disposition) if match: filename = match.group(1).strip('"') if not filename: filename = f'file_{idx}.unknown' save_path = os.path.join('pics', filename) with open(save_path, 'wb') as f: f.write(response.read()) print(f"已保存: {save_path}") except Exception as e: print(f"下载失败 {url}: {str(e)}")
关键细节说明
- 正则表达式兼容两种常见的
Content-Disposition格式:attachment; filename="example.jpg"和attachment; filename=example.pdf。 - 用
enumerate替代links.index(item),避免重复链接导致的索引错误,同时提升效率。 - 加入异常处理,避免单个链接下载失败导致整个脚本中断。
- 提前创建保存目录,防止文件保存时因目录不存在报错。
内容的提问来源于stack exchange,提问作者murd0x
相关产品推荐
相关产品推荐

