You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Curl Cookie认证批量下载archive.org文件遇401未授权问题

问题描述

我希望批量下载archive.org中预定义列表内的文件以提升效率。无需认证的文件可正常下载,但需要登录的文件下载失败。archive.org采用Cookie认证,我尝试了以下命令:

curl --dump-header cookie.txt -u user:pass -H "Connection: keep-alive" https://archive.org/account/login
curl -v -b cookie.txt --location-trusted -O https://archive.org/download/mame-merged/mame-merged/3countb.zip

使用--verbose查看发现重定向的CDN请求中已携带Cookie,但始终返回401未授权:

> Cookie: donation-identifier=3696401840959dfcc5e3b3f02cd51945; abtest-identifier=d6bb7a8a590bb5640d31ce8d27a112ce; test-cookie=1; ia-auth=fbda0754bc3108de49a9e811460d8e59
>
* Request completely sent off
0     0    0     0    0     0      0      0 --:--:--  0:00:03 --:--:--     0< HTTP/1.1 401 Unauthorized
< Server: nginx
< Date: Sun, 26 Jan 2025 15:56:03 GMT
< Content-Length: 172
< Connection: keep-alive

请问我遗漏了什么?

解决方法

你的登录命令逻辑错误:archive.org的登录不是HTTP基本认证,不能用-u参数直接提交账号密码,需要通过表单POST完成登录流程,同时要携带CSRF令牌。另外,你用--dump-header保存Cookie的方式不规范,会导致关键会话Cookie缺失。

正确的curl登录+下载步骤

  1. 获取登录页面的CSRF令牌
    先访问登录页,保存初始Cookie并提取CSRF令牌:

    curl https://archive.org/account/login -c cookie.txt -o login_page.html
    # 提取CSRF令牌(需grep、sed工具)
    CSRF_TOKEN=$(grep 'name="csrfmiddlewaretoken"' login_page.html | sed 's/.*value="\([^"]*\)".*/\1/')
    
  2. 提交登录表单获取有效会话Cookie
    用POST请求提交账号密码和CSRF令牌,更新Cookie文件:

    curl https://archive.org/account/login -b cookie.txt -c cookie.txt \
    -d "csrfmiddlewaretoken=$CSRF_TOKEN" \
    -d "username=你的账号" \
    -d "password=你的密码" \
    -d "submit=Log+In"
    
  3. 使用有效Cookie下载文件
    现在用保存的Cookie文件发起下载请求,跟随重定向:

    curl -b cookie.txt -L -O https://archive.org/download/mame-merged/mame-merged/3countb.zip
    

批量下载更高效的方案

如果是批量操作,推荐使用archive.org官方的internetarchive命令行工具,步骤更简单:

  1. 安装工具:
    pip install internetarchive
    
  2. 配置账号:
    ia configure
    
  3. 批量下载指定文件(可从列表读取):
    # 单文件示例
    ia download mame-merged --files 3countb.zip
    # 批量从文件读取列表,假设文件是file_list.txt,每行一个文件名
    ia download mame-merged --files @file_list.txt
    

关键注意点

  • 不要用--dump-header保存Cookie,改用-c(存Cookie到文件)和-b(读Cookie),这是curl处理Cookie的标准方式。
  • archive.org的CDN需要完整的会话Cookie(如ia-user、ia-session),之前的请求只拿到了ia-auth,缺少这些关键标识导致401。

内容的提问来源于stack exchange,提问作者silverchair

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 18:13:26