You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取无显式真实链接的HTML a标签下载链接并使用Python批量下载

实现方案

第一步:抓包确认下载接口规则

先打开浏览器开发者工具的「网络」面板,点击一次下载按钮,找到对应的下载请求,记录:

  • 请求方式(GET/POST)
  • 请求URL的固定前缀
  • 参数传递方式(URL查询参数/请求体参数)
    这类动态下载接口大概率为GET请求,格式通常类似:https://[你的站点域名]/xxx/download?clazzId={clazzId}&libraryId={libraryId}&relationId={relationId}

第二步:Python批量实现代码

首先安装依赖库:

pip install requests beautifulsoup4

代码示例:

import requests
import json
import os
import time
from bs4 import BeautifulSoup

# 配置项,请替换为你的实际参数
# 站点根域名
BASE_URL = "https://你的站点域名.com"
# 下载接口模板,替换为你抓包得到的实际接口格式
DOWNLOAD_API_TPL = "https://你的站点域名.com/api/download?clazzId={clazzId}&libraryId={libraryId}&relationId={relationId}"
# 请求头,Cookie从浏览器已登录的请求头中复制
HEADERS = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36",
    "Cookie": "你的登录Cookie字符串"
}
# 文件保存目录
SAVE_DIR = "./download_files"
os.makedirs(SAVE_DIR, exist_ok=True)

# 1. 获取列表页HTML源码
# 方式1:直接请求列表页
list_page_url = "下载列表页的实际URL"
resp = requests.get(list_page_url, headers=HEADERS)
resp.encoding = "utf-8"
html_content = resp.text
# 方式2:读取本地保存的HTML文件
# with open("list_page.html", "r", encoding="utf-8") as f:
#     html_content = f.read()

# 2. 解析HTML提取下载参数
soup = BeautifulSoup(html_content, "lxml")
# 匹配所有下载按钮的a标签
download_tags = soup.select("a.download_ic.checkSafe")

for index, tag in enumerate(download_tags):
    try:
        # 提取data属性里的参数
        data_info = json.loads(tag.get("data"))
        # 拼接下载链接
        download_url = DOWNLOAD_API_TPL.format(
            clazzId=data_info["clazzId"],
            libraryId=data_info["libraryId"],
            relationId=data_info["relationId"]
        )
        # 提取文件名(取第一个td的title属性)
        tr_parent = tag.find_parent("tr")
        file_name = tr_parent.find("td").get("title", f"file_{index}") + ".zip"
        save_path = os.path.join(SAVE_DIR, file_name)

        # 3. 下载文件
        print(f"正在下载:{file_name}")
        file_resp = requests.get(download_url, headers=HEADERS, timeout=120, allow_redirects=True)
        with open(save_path, "wb") as f:
            f.write(file_resp.content)
        print(f"下载完成:{save_path}")
        # 控制请求频率,避免触发反爬
        time.sleep(1)
    except Exception as e:
        print(f"第{index+1}个文件下载失败:{str(e)}")
        continue

注意事项

  • 如果站点反爬严格,可适当延长time.sleep的等待时长
  • Cookie过期后需要重新从浏览器中复制替换
  • 若下载接口为POST请求,将requests.get改为requests.post,对应参数放到json或data入参中即可
  • 若下载链接需要从302跳转中获取,保持allow_redirects=True即可自动跳转拿到文件内容

内容的提问来源于stack exchange,提问作者user15964

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 05:12:02