使用wget/curl下载m3u文件遇404错误,浏览器刷新可正常下载
问题描述
需要下载一个m3u文件,在Firefox中首次访问目标URL时出现404错误,但刷新页面后下载可正常启动。使用Python调用curl脚本下载时也遇到同样的404问题,相关代码及错误页面如下:
代码片段
import os url = 'https://watch.xpsclub.com:2083/get.php?username=FOO&password=XXX&output=ts&type=m3u_plus' M3UPATH= '/tmp/IPTVWORLD55.m3u' def create_bouquet(): if not os.path.exists(M3UPATH): #os.system('wget --no-check-certificate -q -O- --trust-server-names %s > %s' % (url, M3UPATH)) #os.system('curl -c --limit-rate 50K %s -o %s' % (url, M3UPATH)) os.system('curl -H "Accept-Charset: utf-8" -H "Content-Type: application/x-www-form-urlencoded" --limit-rate 100K %s -o %s' % (url, M3UPATH)) create_bouquet()
404错误页面
<html> <head><title>404 Not Found</title></head> <body> <center><h1>404 Not Found</h1></center> <hr><center>nginx</center> </body> </html> <!-- a padding to disable MSIE and Chrome friendly error page --> <!-- a padding to disable MSIE and Chrome friendly error page --> <!-- a padding to disable MSIE and Chrome friendly error page --> <!-- a padding to disable MSIE and Chrome friendly error page --> <!-- a padding to disable MSIE and Chrome friendly error page --> <!-- a padding to disable MSIE and Chrome friendly error page -->
解决方案
1. 模拟浏览器完整请求头
服务器大概率通过校验请求头判断是否为合法请求,curl默认请求头与浏览器差异大,导致首次请求被拦截。可以添加浏览器真实请求头,并启用Cookie复用(刷新时浏览器会自动携带之前的Cookie):
os.system('curl -c /tmp/cookies.txt -b /tmp/cookies.txt -H "User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:109.0) Gecko/20100101 Firefox/117.0" -H "Accept: text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8" -H "Accept-Language: zh-CN,zh;q=0.8,en-US;q=0.5,en;q=0.3" -H "Connection: keep-alive" --limit-rate 100K %s -o %s' % (url, M3UPATH))
参数说明:
-c /tmp/cookies.txt:保存服务器返回的Cookie到本地文件-b /tmp/cookies.txt:发送请求时携带已保存的Cookie- 复制Firefox开发者工具中真实请求的头信息,让请求行为更贴近浏览器
2. 添加请求重试逻辑
既然刷新可解决问题,可在脚本中增加重试机制,首次请求失败后自动重试:
import os import subprocess url = 'https://watch.xpsclub.com:2083/get.php?username=FOO&password=XXX&output=ts&type=m3u_plus' M3UPATH= '/tmp/IPTVWORLD55.m3u' def create_bouquet(): if not os.path.exists(M3UPATH): # 首次请求 result = subprocess.run( ['curl', '-H', 'Accept-Charset: utf-8', '--limit-rate', '100K', url, '-o', M3UPATH], capture_output=True ) # 检查请求是否失败(返回码非0或文件内容为404页面) if result.returncode != 0: # 携带Cookie重试 subprocess.run( ['curl', '-c', '/tmp/cookies.txt', '-b', '/tmp/cookies.txt', '-H', 'User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:109.0) Gecko/20100101 Firefox/117.0', '--limit-rate', '100K', url, '-o', M3UPATH] ) create_bouquet()
用subprocess替代os.system,能更精准判断请求状态,触发重试逻辑。
3. 确认URL参数有效性
检查URL中的参数格式,确保&等符号未被错误编码(你的代码中已正确使用&而非&,这部分无需调整)。
原因分析
这种情况基本是服务器的会话验证/反爬机制导致:
- 首次请求时,缺少合法会话标识(如Cookie)或请求头不符合浏览器特征,被服务器拦截返回404
- 刷新时浏览器自动携带了服务器之前设置的Cookie,且请求头更完整,通过服务器校验
- 部分服务器的CDN节点需要首次请求触发资源预热,第二次请求即可命中缓存返回正常内容
内容的提问来源于stack exchange,提问作者mino31
相关产品推荐
相关产品推荐

