You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Botify API POST请求失败求助(错误码1020)

Botify导出任务创建失败求助

我按照Botify官方文档发送POST请求创建导出任务,但任务持续失败,错误提示未知集合。作为Botify新手,详细信息如下:

Python代码

import requests
import json

token = 'XXXXXXXXXXXXXXXXXXXXXXX'

url = "https://api.botify.com/v1/jobs"

json_payload = {
  "job_type": "export",
  "payload": {
    "username": "XYZ0",
    "project": "XYZ.com",
    "export_size": 50,
    "formatter": "csv",
    "formatter_config": {
            "delimiter": ",",
            "print_delimiter": False,
            "print_header": True,
            "header_format": "verbose"
        },
    "connector": "direct_download",
    "extra_config": {},
    "query": {
      "collections": ["crawl.20230509"],
      "query": {
        "dimensions": ["url",
                       "crawl.20230509.date_crawled",
                       "crawl.20230509.content_type",
                       "crawl.20230509.http_code"
                       ],
        "metrics": [],
        "sort": [1]
      }
    }
  }
}

headers = {
    "Content-Type": "application/json",
    "Authorization": f"Token {token}"
}

response = requests.post(url, json=json_payload, headers=headers).json()

print(response)

注:不确定extra_config字段内容,传入了空字典。

错误信息

{'status': 400, 'error': {'error_code': '1020', 'message': 'Badly formatted request', 'error_detail': {'payload': {'query': {'collections': ['Unknown collection "crawl.20230509".']}}}}}

补充说明

以下是一段可正常运行的相似示例代码,但我不清楚自己的代码问题出在哪里:

url = "https://api.botify.com/v1/jobs"

token = 'XXXXXXXXXXXXXXXXXXXXXXXX'

headers = {
    "Content-Type": "application/json",
    "Authorization": f"Token {token}"    
}

data = """
{
  "job_type": "export",
  "payload": {
    "username": "XXXX0",
    "project": "XXXXX.com",
    "connector": "direct_download",
    "formatter": "csv",
    "formatter_config": {
            "delimiter": ",",
            "print_delimiter": false,
            "print_header": true,
            "header_format": "verbose"
        },
    "export_size": 50,
    "query": {
      "collections": ["crawl.20250515"],
      "query": {
        "dimensions": ["url",
"crawl.20221012.date_crawled",
"crawl.20221012.content_type",
"crawl.20221012.http_code"
],
        "metrics": [],
        "sort": [1]
      }
"""

问题排查与解决方案

核心问题

错误提示明确指出crawl.20230509是未知集合,说明你指定的爬取任务ID不存在于你的项目中。

解决步骤

  1. 获取有效爬取任务ID
    调用Botify爬取列表接口,获取当前项目下所有真实存在的爬取任务ID:

    import requests
    
    token = '你的token'
    username = 'XYZ0'
    project = 'XYZ.com'
    url = f"https://api.botify.com/v1/users/{username}/projects/{project}/crawls"
    
    headers = {
        "Authorization": f"Token {token}"
    }
    
    response = requests.get(url, headers=headers).json()
    print(response)
    

    返回结果中会列出格式为crawl.YYYYMMDD的有效爬取ID。

  2. 替换代码中的无效ID
    将代码中collections数组和dimensions字段里的crawl.20230509,全部替换为上一步获取到的有效爬取ID。

  3. 额外说明

    • extra_config字段为空字典是允许的,不是当前错误的诱因。
    • 示例代码中collections与dimensions使用不同爬取ID属于笔误,你的代码中保持ID一致是正确做法。

内容的提问来源于stack exchange,提问作者marie20

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 17:47:17