调用Scilit的articles接口遇400缺失参数错误,如何解决?
问题描述
要抓取Scilit网站(目标页面:https://www.scilit.net/publications?facet=%7B%22pubyear%22%3A%5B2024%5D%2C%22language%22%3A%5B%22English%22%5D%2C%22publication_type%22%3A%5B%22JOURNAL-ARTICLE%22%5D%7D)Network面板中名为"articles"的XHR请求数据,编写了如下Python代码:
import cloudscraper import json url = "https://www.scilit.net/api/solr/articles" headers = { "user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36", "Content-Type": "application/json" } payload = { "bPublic": True, "pub_year": 2024, "language": "English", "type": "JOURNAL-ARTICLE" } scraper = cloudscraper.create_scraper() response = scraper.post(url, params=headers) print(response.status_code) print(response.content) try: json_data = response.json() file_path = "data.json" with open(file_path, "w") as json_file: json.dump(json_data, json_file, indent=4) print(f'JSON data successfully saved to {file_path}') except ValueError: print('Response is not valid JSON.') json_data = response.text else: print(f'Request failed with status code {response.status_code}') file_path = "json.data"
运行后出现以下问题:
- 返回200状态码,但响应内容为
{"error":400,"error_msg":"Missing parameter"} - 同时打印出
JSON data successfully saved to data.json和Request failed with status code 200 - 服务器提示缺失参数,但无法确定具体缺少哪个参数
已尝试操作:
- 确保headers与浏览器一致
- 确认payload包含
pub_year、language、type等参数 - 使用
cloudscraper绕过反爬
错误原因
- 请求参数传递错误:代码中调用
scraper.post()时,把请求头通过params参数传递(这会将请求头转为URL查询参数),而定义好的payload根本没传给服务器,这是触发"Missing parameter"的核心原因。 - 异常处理逻辑混乱:
try-except-else的逻辑使用错误,else块会在try块无异常时执行,所以即使返回错误JSON,也会执行"请求失败"的打印,导致输出混乱。 - payload参数不全:浏览器实际发送的XHR请求中,除已定义的参数外,还包含分页(
start、rows)、排序(sort)等必填参数,这些参数被遗漏。
修复方案
1. 修正请求参数传递方式
调用post方法时,通过headers参数传请求头,通过json参数传payload(服务器接收JSON格式请求体)。
2. 重构异常处理逻辑
先判断请求状态码,再解析响应,避免逻辑冲突;增加错误JSON的判断,区分正常数据与服务器错误。
3. 补充完整payload参数
对照浏览器Network面板中"articles"请求的Payload,补充所有必填参数,比如分页、排序、搜索关键词等。
修正后的代码
import cloudscraper import json url = "https://www.scilit.net/api/solr/articles" headers = { "user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36", "Content-Type": "application/json", "Referer": "https://www.scilit.net/publications?facet=%7B%22pubyear%22%3A%5B2024%5D%2C%22language%22%3A%5B%22English%22%5D%2C%22publication_type%22%3A%5B%22JOURNAL-ARTICLE%22%5D%7D", "X-Requested-With": "XMLHttpRequest" } # 补充浏览器实际发送的完整参数 payload = { "bPublic": True, "pub_year": 2024, "language": "English", "type": "JOURNAL-ARTICLE", "start": 0, "rows": 20, "sort": "pubyear desc,score desc", "q": "*" } scraper = cloudscraper.create_scraper() response = scraper.post(url, headers=headers, json=payload) print(f"状态码: {response.status_code}") # 先判断请求是否成功 if response.status_code == 200: try: json_data = response.json() # 检查返回的是否是错误数据 if "error" in json_data: print(f"服务器错误: {json_data['error_msg']}") else: file_path = "data.json" with open(file_path, "w", encoding="utf-8") as json_file: json.dump(json_data, json_file, indent=4, ensure_ascii=False) print(f'JSON数据已成功保存到 {file_path}') except ValueError: print('响应不是有效的JSON格式') with open("response.txt", "w", encoding="utf-8") as f: f.write(response.text) print('响应内容已保存到response.txt') else: print(f'请求失败,状态码: {response.status_code}') with open("error_response.txt", "w", encoding="utf-8") as f: f.write(response.text)
额外说明
- 可以在浏览器Network面板中,查看"articles"请求的Payload和Headers,确保代码中的参数和请求头与浏览器完全一致,若有动态生成参数需同步更新。
- 如果仍被反爬拦截,可尝试添加从浏览器复制的
Cookie参数(注意Cookie可能过期,需定期更新)。
内容的提问来源于stack exchange,提问作者Mohsin ali
相关产品推荐
相关产品推荐

