You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

调用Scilit的articles接口遇400缺失参数错误,如何解决?

问题描述

要抓取Scilit网站(目标页面:https://www.scilit.net/publications?facet=%7B%22pubyear%22%3A%5B2024%5D%2C%22language%22%3A%5B%22English%22%5D%2C%22publication_type%22%3A%5B%22JOURNAL-ARTICLE%22%5D%7D)Network面板中名为"articles"的XHR请求数据,编写了如下Python代码:

import cloudscraper
import json

url = "https://www.scilit.net/api/solr/articles"

headers = {
    "user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36",
    "Content-Type": "application/json"
}

payload = {
    "bPublic": True,
    "pub_year": 2024,
    "language": "English",
    "type": "JOURNAL-ARTICLE"
}

scraper = cloudscraper.create_scraper()
response = scraper.post(url, params=headers)
print(response.status_code)
print(response.content)

try:
    json_data = response.json()
    file_path = "data.json"

    with open(file_path, "w") as json_file:
        json.dump(json_data, json_file, indent=4)
        print(f'JSON data successfully saved to {file_path}')

except ValueError:
    print('Response is not valid JSON.')
    json_data = response.text

else:
    print(f'Request failed with status code {response.status_code}')
    file_path = "json.data"

运行后出现以下问题:

  • 返回200状态码,但响应内容为{"error":400,"error_msg":"Missing parameter"}
  • 同时打印出JSON data successfully saved to data.json和Request failed with status code 200
  • 服务器提示缺失参数,但无法确定具体缺少哪个参数

已尝试操作:

  • 确保headers与浏览器一致
  • 确认payload包含pub_year、language、type等参数
  • 使用cloudscraper绕过反爬

错误原因

  1. 请求参数传递错误:代码中调用scraper.post()时,把请求头通过params参数传递(这会将请求头转为URL查询参数),而定义好的payload根本没传给服务器,这是触发"Missing parameter"的核心原因。
  2. 异常处理逻辑混乱:try-except-else的逻辑使用错误,else块会在try块无异常时执行,所以即使返回错误JSON,也会执行"请求失败"的打印,导致输出混乱。
  3. payload参数不全:浏览器实际发送的XHR请求中,除已定义的参数外,还包含分页(start、rows)、排序(sort)等必填参数,这些参数被遗漏。

修复方案

1. 修正请求参数传递方式

调用post方法时,通过headers参数传请求头,通过json参数传payload(服务器接收JSON格式请求体)。

2. 重构异常处理逻辑

先判断请求状态码,再解析响应,避免逻辑冲突;增加错误JSON的判断,区分正常数据与服务器错误。

3. 补充完整payload参数

对照浏览器Network面板中"articles"请求的Payload,补充所有必填参数,比如分页、排序、搜索关键词等。

修正后的代码

import cloudscraper
import json

url = "https://www.scilit.net/api/solr/articles"

headers = {
    "user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36",
    "Content-Type": "application/json",
    "Referer": "https://www.scilit.net/publications?facet=%7B%22pubyear%22%3A%5B2024%5D%2C%22language%22%3A%5B%22English%22%5D%2C%22publication_type%22%3A%5B%22JOURNAL-ARTICLE%22%5D%7D",
    "X-Requested-With": "XMLHttpRequest"
}

# 补充浏览器实际发送的完整参数
payload = {
    "bPublic": True,
    "pub_year": 2024,
    "language": "English",
    "type": "JOURNAL-ARTICLE",
    "start": 0,
    "rows": 20,
    "sort": "pubyear desc,score desc",
    "q": "*"
}

scraper = cloudscraper.create_scraper()
response = scraper.post(url, headers=headers, json=payload)

print(f"状态码: {response.status_code}")

# 先判断请求是否成功
if response.status_code == 200:
    try:
        json_data = response.json()
        # 检查返回的是否是错误数据
        if "error" in json_data:
            print(f"服务器错误: {json_data['error_msg']}")
        else:
            file_path = "data.json"
            with open(file_path, "w", encoding="utf-8") as json_file:
                json.dump(json_data, json_file, indent=4, ensure_ascii=False)
            print(f'JSON数据已成功保存到 {file_path}')
    except ValueError:
        print('响应不是有效的JSON格式')
        with open("response.txt", "w", encoding="utf-8") as f:
            f.write(response.text)
        print('响应内容已保存到response.txt')
else:
    print(f'请求失败,状态码: {response.status_code}')
    with open("error_response.txt", "w", encoding="utf-8") as f:
        f.write(response.text)

额外说明

  • 可以在浏览器Network面板中,查看"articles"请求的Payload和Headers,确保代码中的参数和请求头与浏览器完全一致,若有动态生成参数需同步更新。
  • 如果仍被反爬拦截,可尝试添加从浏览器复制的Cookie参数(注意Cookie可能过期,需定期更新)。

内容的提问来源于stack exchange,提问作者Mohsin ali

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 08:52:33