如何获取无显式真实链接的HTML a标签下载链接并使用Python批量下载
实现方案
第一步:抓包确认下载接口规则
先打开浏览器开发者工具的「网络」面板,点击一次下载按钮,找到对应的下载请求,记录:
- 请求方式(GET/POST)
- 请求URL的固定前缀
- 参数传递方式(URL查询参数/请求体参数)
这类动态下载接口大概率为GET请求,格式通常类似:https://[你的站点域名]/xxx/download?clazzId={clazzId}&libraryId={libraryId}&relationId={relationId}
第二步:Python批量实现代码
首先安装依赖库:
pip install requests beautifulsoup4
代码示例:
import requests import json import os import time from bs4 import BeautifulSoup # 配置项,请替换为你的实际参数 # 站点根域名 BASE_URL = "https://你的站点域名.com" # 下载接口模板,替换为你抓包得到的实际接口格式 DOWNLOAD_API_TPL = "https://你的站点域名.com/api/download?clazzId={clazzId}&libraryId={libraryId}&relationId={relationId}" # 请求头,Cookie从浏览器已登录的请求头中复制 HEADERS = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36", "Cookie": "你的登录Cookie字符串" } # 文件保存目录 SAVE_DIR = "./download_files" os.makedirs(SAVE_DIR, exist_ok=True) # 1. 获取列表页HTML源码 # 方式1:直接请求列表页 list_page_url = "下载列表页的实际URL" resp = requests.get(list_page_url, headers=HEADERS) resp.encoding = "utf-8" html_content = resp.text # 方式2:读取本地保存的HTML文件 # with open("list_page.html", "r", encoding="utf-8") as f: # html_content = f.read() # 2. 解析HTML提取下载参数 soup = BeautifulSoup(html_content, "lxml") # 匹配所有下载按钮的a标签 download_tags = soup.select("a.download_ic.checkSafe") for index, tag in enumerate(download_tags): try: # 提取data属性里的参数 data_info = json.loads(tag.get("data")) # 拼接下载链接 download_url = DOWNLOAD_API_TPL.format( clazzId=data_info["clazzId"], libraryId=data_info["libraryId"], relationId=data_info["relationId"] ) # 提取文件名(取第一个td的title属性) tr_parent = tag.find_parent("tr") file_name = tr_parent.find("td").get("title", f"file_{index}") + ".zip" save_path = os.path.join(SAVE_DIR, file_name) # 3. 下载文件 print(f"正在下载:{file_name}") file_resp = requests.get(download_url, headers=HEADERS, timeout=120, allow_redirects=True) with open(save_path, "wb") as f: f.write(file_resp.content) print(f"下载完成:{save_path}") # 控制请求频率,避免触发反爬 time.sleep(1) except Exception as e: print(f"第{index+1}个文件下载失败:{str(e)}") continue
注意事项
- 如果站点反爬严格,可适当延长
time.sleep的等待时长 - Cookie过期后需要重新从浏览器中复制替换
- 若下载接口为POST请求,将
requests.get改为requests.post,对应参数放到json或data入参中即可 - 若下载链接需要从302跳转中获取,保持
allow_redirects=True即可自动跳转拿到文件内容
内容的提问来源于stack exchange,提问作者user15964
相关产品推荐
相关产品推荐

