Botify API POST请求失败求助(错误码1020)
Botify导出任务创建失败求助
我按照Botify官方文档发送POST请求创建导出任务,但任务持续失败,错误提示未知集合。作为Botify新手,详细信息如下:
Python代码
import requests import json token = 'XXXXXXXXXXXXXXXXXXXXXXX' url = "https://api.botify.com/v1/jobs" json_payload = { "job_type": "export", "payload": { "username": "XYZ0", "project": "XYZ.com", "export_size": 50, "formatter": "csv", "formatter_config": { "delimiter": ",", "print_delimiter": False, "print_header": True, "header_format": "verbose" }, "connector": "direct_download", "extra_config": {}, "query": { "collections": ["crawl.20230509"], "query": { "dimensions": ["url", "crawl.20230509.date_crawled", "crawl.20230509.content_type", "crawl.20230509.http_code" ], "metrics": [], "sort": [1] } } } } headers = { "Content-Type": "application/json", "Authorization": f"Token {token}" } response = requests.post(url, json=json_payload, headers=headers).json() print(response)
注:不确定extra_config字段内容,传入了空字典。
错误信息
{'status': 400, 'error': {'error_code': '1020', 'message': 'Badly formatted request', 'error_detail': {'payload': {'query': {'collections': ['Unknown collection "crawl.20230509".']}}}}}
补充说明
以下是一段可正常运行的相似示例代码,但我不清楚自己的代码问题出在哪里:
url = "https://api.botify.com/v1/jobs" token = 'XXXXXXXXXXXXXXXXXXXXXXXX' headers = { "Content-Type": "application/json", "Authorization": f"Token {token}" } data = """ { "job_type": "export", "payload": { "username": "XXXX0", "project": "XXXXX.com", "connector": "direct_download", "formatter": "csv", "formatter_config": { "delimiter": ",", "print_delimiter": false, "print_header": true, "header_format": "verbose" }, "export_size": 50, "query": { "collections": ["crawl.20250515"], "query": { "dimensions": ["url", "crawl.20221012.date_crawled", "crawl.20221012.content_type", "crawl.20221012.http_code" ], "metrics": [], "sort": [1] } """
问题排查与解决方案
核心问题
错误提示明确指出crawl.20230509是未知集合,说明你指定的爬取任务ID不存在于你的项目中。
解决步骤
获取有效爬取任务ID
调用Botify爬取列表接口,获取当前项目下所有真实存在的爬取任务ID:import requests token = '你的token' username = 'XYZ0' project = 'XYZ.com' url = f"https://api.botify.com/v1/users/{username}/projects/{project}/crawls" headers = { "Authorization": f"Token {token}" } response = requests.get(url, headers=headers).json() print(response)返回结果中会列出格式为
crawl.YYYYMMDD的有效爬取ID。替换代码中的无效ID
将代码中collections数组和dimensions字段里的crawl.20230509,全部替换为上一步获取到的有效爬取ID。额外说明
extra_config字段为空字典是允许的,不是当前错误的诱因。- 示例代码中
collections与dimensions使用不同爬取ID属于笔误,你的代码中保持ID一致是正确做法。
内容的提问来源于stack exchange,提问作者marie20
相关产品推荐
相关产品推荐

