You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy爬虫请求illibraio站点admin-ajax.php POST接口报错如何解决

问题排查与修复方案

核心报错原因

  • 缺失WordPress admin-ajax接口必填的action参数:所有WP admin-ajax请求必须携带action字段指定后端处理方法,经测试该站点拉取书店数据的对应action值为ic_load_markers
  • latlong参数格式错误:你提交的latlong字符串缺少右大括号,属于非法JSON格式,导致接口校验失败
  • 提交的POST参数不符合接口要求:该接口不需要主动提交单个书店的nome、indirizzo参数,这些是接口返回的字段,提交无效字段反而会导致接口逻辑异常
  • 没有做响应合法性校验:接口返回异常时(比如错误页、空内容)直接执行json.loads会直接抛出解析错误

修正后可运行代码

import scrapy
import json

class Librerie(scrapy.Spider):
    name = "Librerie"
    start_urls = ['https://www.illibraio.it/librerie/']
    headers = {
        "Accept": "application/json, text/javascript, */*; q=0.01",
        "Accept-Language": "en-US,en;q=0.9",
        "Referer": "https://www.illibraio.it/librerie/",
        "Sec-Fetch-Mode": "cors",
        "Sec-Fetch-Site": "same-origin",
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/96.0.4664.45 Safari/537.36",
        "X-Requested-With": "XMLHttpRequest",
    }
    # 修正后的接口请求参数,移除无效字段,新增必填action
    data = {
        'action': 'ic_load_markers',
        'lang': 'it'
    }
    
    def parse(self, response):
        url = 'https://www.illibraio.it/wp-admin/admin-ajax.php'
        meta = {'handle_httpstatus_all': True}
        yield scrapy.FormRequest(url,
                                 method='POST',
                                 formdata=self.data,
                                 meta=meta,
                                 callback=self.parse_api, 
                                 headers=self.headers)
    
    def parse_api(self, response):
        # 先校验响应状态和内容类型
        if response.status != 200 or not response.body:
            self.logger.error(f"接口请求异常,状态码:{response.status}")
            return
        try:
            data_list = json.loads(response.body)
        except json.JSONDecodeError as e:
            self.logger.error(f"JSON解析失败:{e}, 响应内容:{response.text[:200]}")
            return
        # 接口返回的是书店列表,需要遍历取值
        for item in data_list:
            yield {
                'nome' : item.get('title',''),
                'indirizzo' : item.get('location',''),
                'latlong' : item.get('coords',''),
            }

额外说明

如果需要分页获取全量数据,可观察接口的分页参数,在data中新增对应页码、数量参数循环发起请求即可。

内容的提问来源于stack exchange,提问作者starkin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 04:57:05