You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用FormRequest通过HTTP POST提取数据遇问题求助

Scrapy爬取bitsclassic.com无输出问题排查

核心问题分析

你的代码存在几个关键错误,导致无输出:

  • 接口响应处理逻辑错误:parse_page 直接处理 Product/ProductList 接口的响应,但这个接口返回的是产品列表数据(不是单个产品详情页),你用产品详情页的选择器自然拿不到任何数据。
  • 请求参数可能不全:该AJAX接口可能需要额外请求头或参数(比如X-Requested-With: XMLHttpRequest标识AJAX请求,或者Page参数),仅传Cats和Size可能无法触发正确响应。
  • category_id提取无异常处理:如果分类URL格式不符合正则/(\d+)-,会直接抛出异常中断爬虫,导致后续请求无法执行。
  • 未从列表接口提取产品URL:你跳过了从列表接口获取单个产品链接的步骤,直接尝试从列表响应解析详情数据,逻辑完全错误。

修正后的代码示例

import re
import scrapy
from scrapy import FormRequest
from your_project_name.items import BitsclassicItem  # 替换为你的项目Item路径

class BitsclassicSpider(scrapy.Spider):
    name = "bitsclassic"
    start_urls = ['https://bitsclassic.com/fa']

    def parse(self, response):
        # 提取分类URL,跳过第一个可能的无效链接
        category_urls = response.css('ul.children a::attr(href)').getall()[1:]
        for category_url in category_urls:
            yield scrapy.Request(category_url, callback=self.parse_category)

    def parse_category(self, response):
        # 提取分类ID,增加异常处理
        category_match = re.search(r"/(\d+)-", response.url)
        if not category_match:
            self.logger.warning(f"无法从URL提取分类ID: {response.url}")
            return
        
        category_id = category_match.group(1)
        form_data = {
            'Cats': category_id,
            'Size': '1000',
            'Page': '1'  # 补充Page参数,部分接口需要
        }

        yield FormRequest(
            url='https://bitsclassic.com/fa/Product/ProductList',
            method='POST',
            formdata=form_data,
            headers={
                # 补充AJAX请求头,模拟前端请求
                'X-Requested-With': 'XMLHttpRequest',
                'Referer': response.url
            },
            callback=self.parse_product_list
        )

    def parse_product_list(self, response):
        # 从列表接口响应中提取所有产品详情页URL
        product_urls = response.css('div.product-item a::attr(href)').getall()
        for product_url in product_urls:
            # 拼接完整URL(如果是相对路径)
            full_product_url = response.urljoin(product_url)
            yield scrapy.Request(full_product_url, callback=self.parse_product_detail)

    def parse_product_detail(self, response):
        # 解析单个产品详情页
        title = response.css('p[itemrolep="name"]::text').get()
        price = response.xpath('//div[@id="priceBox"]//span[@data-role="price"]/text()').get()
        categories = response.xpath('//div[@class="con-main"]//a/text()').getall()

        product_exist = True
        if price:
            price = price.strip()
        else:
            price = None
            product_exist = False

        item = BitsclassicItem()
        item["title"] = title.strip() if title else None
        item["categories"] = categories[3:-1] if len(categories) >=4 else []
        item["product_exist"] = product_exist
        item["price"] = price
        item["url"] = response.url
        item["domain"] = "bitsclassic.com/fa"

        yield item

额外排查建议

  1. 用浏览器开发者工具(F12)查看Product/ProductList接口的实际请求参数和响应格式,确认是否有遗漏的参数(比如Sort、Order等)。
  2. 启用Scrapy日志(LOG_LEVEL=DEBUG),查看请求的状态码、响应内容,确认接口是否返回预期数据。
  3. 检查是否有反爬机制(比如Cookie验证),如果需要,可以在请求中携带从start_url获取的Cookie。

内容的提问来源于stack exchange,提问作者Amir Hamidi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 09:27:45