如何用Python Requests处理HTML Meta重定向以获取目标文件?
解决requests无法处理meta refresh重定向的问题
问题原因
requests库的allow_redirects=True只支持HTTP协议级别的3xx状态码重定向,而你用的<meta http-equiv="refresh">是浏览器端的页面跳转逻辑,requests不会自动解析HTML并执行这个跳转——所以你之前的代码实际上是把包含meta标签的HTML内容存成了zip文件,自然不是目标文件。
解决方案
方法1:手动解析HTML提取重定向URL(轻量推荐)
用BeautifulSoup解析HTML页面,提取meta refresh中的目标URL,再单独请求该URL获取文件:
首先安装依赖:
pip install beautifulsoup4 requests
代码示例:
import requests from bs4 import BeautifulSoup # 初始重定向页面URL redirect_page_url = 'https://gi-b716.github.io/modrinth-api-tools/api/version/d/lasted' response = requests.get(redirect_page_url) # 解析HTML找到meta refresh标签 soup = BeautifulSoup(response.text, 'html.parser') refresh_meta = soup.find('meta', attrs={'http-equiv': 'refresh'}) if refresh_meta: # 拆分content属性获取目标URL content_str = refresh_meta.get('content', '') if 'url=' in content_str: target_url = content_str.split('url=')[1].strip() # 请求目标文件 file_response = requests.get(target_url) with open("1.2.0.zip", "wb") as f: f.write(file_response.content) print("文件下载完成") else: print("未找到meta重定向标签")
如果你的HTML结构非常简单(只有那一行meta标签),也可以不用BeautifulSoup,直接用字符串匹配提取URL,省去依赖:
import requests redirect_page_url = 'https://gi-b716.github.io/modrinth-api-tools/api/version/d/lasted' response = requests.get(redirect_page_url) # 简单匹配提取重定向URL text_content = response.text if 'meta http-equiv="refresh"' in text_content: start_idx = text_content.find('url=') + 4 end_idx = text_content.find('"', start_idx) target_url = text_content[start_idx:end_idx] file_response = requests.get(target_url) with open("1.2.0.zip", "wb") as f: f.write(file_response.content)
方法2:用模拟浏览器工具(重但更贴近真实浏览器行为)
如果需要完全模拟浏览器的所有行为(比如处理JS跳转、Cookie等),可以用selenium或playwright,但需要额外安装浏览器驱动,适合复杂场景:
以selenium为例:
pip install selenium
代码示例(需提前下载对应浏览器的驱动,比如ChromeDriver):
from selenium import webdriver import time import requests driver = webdriver.Chrome() driver.get('https://gi-b716.github.io/modrinth-api-tools/api/version/d/lasted') # 等待页面跳转完成 time.sleep(1) # 获取当前页面的URL(即跳转后的文件URL) target_url = driver.current_url driver.quit() # 下载目标文件 response = requests.get(target_url) with open("1.2.0.zip", "wb") as f: f.write(response.content)
内容的提问来源于stack exchange,提问作者GavinCQTD
相关产品推荐
相关产品推荐

