使用FormRequest通过HTTP POST提取数据遇问题求助
Scrapy爬取bitsclassic.com无输出问题排查
核心问题分析
你的代码存在几个关键错误,导致无输出:
- 接口响应处理逻辑错误:
parse_page直接处理Product/ProductList接口的响应,但这个接口返回的是产品列表数据(不是单个产品详情页),你用产品详情页的选择器自然拿不到任何数据。 - 请求参数可能不全:该AJAX接口可能需要额外请求头或参数(比如
X-Requested-With: XMLHttpRequest标识AJAX请求,或者Page参数),仅传Cats和Size可能无法触发正确响应。 - category_id提取无异常处理:如果分类URL格式不符合正则
/(\d+)-,会直接抛出异常中断爬虫,导致后续请求无法执行。 - 未从列表接口提取产品URL:你跳过了从列表接口获取单个产品链接的步骤,直接尝试从列表响应解析详情数据,逻辑完全错误。
修正后的代码示例
import re import scrapy from scrapy import FormRequest from your_project_name.items import BitsclassicItem # 替换为你的项目Item路径 class BitsclassicSpider(scrapy.Spider): name = "bitsclassic" start_urls = ['https://bitsclassic.com/fa'] def parse(self, response): # 提取分类URL,跳过第一个可能的无效链接 category_urls = response.css('ul.children a::attr(href)').getall()[1:] for category_url in category_urls: yield scrapy.Request(category_url, callback=self.parse_category) def parse_category(self, response): # 提取分类ID,增加异常处理 category_match = re.search(r"/(\d+)-", response.url) if not category_match: self.logger.warning(f"无法从URL提取分类ID: {response.url}") return category_id = category_match.group(1) form_data = { 'Cats': category_id, 'Size': '1000', 'Page': '1' # 补充Page参数,部分接口需要 } yield FormRequest( url='https://bitsclassic.com/fa/Product/ProductList', method='POST', formdata=form_data, headers={ # 补充AJAX请求头,模拟前端请求 'X-Requested-With': 'XMLHttpRequest', 'Referer': response.url }, callback=self.parse_product_list ) def parse_product_list(self, response): # 从列表接口响应中提取所有产品详情页URL product_urls = response.css('div.product-item a::attr(href)').getall() for product_url in product_urls: # 拼接完整URL(如果是相对路径) full_product_url = response.urljoin(product_url) yield scrapy.Request(full_product_url, callback=self.parse_product_detail) def parse_product_detail(self, response): # 解析单个产品详情页 title = response.css('p[itemrolep="name"]::text').get() price = response.xpath('//div[@id="priceBox"]//span[@data-role="price"]/text()').get() categories = response.xpath('//div[@class="con-main"]//a/text()').getall() product_exist = True if price: price = price.strip() else: price = None product_exist = False item = BitsclassicItem() item["title"] = title.strip() if title else None item["categories"] = categories[3:-1] if len(categories) >=4 else [] item["product_exist"] = product_exist item["price"] = price item["url"] = response.url item["domain"] = "bitsclassic.com/fa" yield item
额外排查建议
- 用浏览器开发者工具(F12)查看
Product/ProductList接口的实际请求参数和响应格式,确认是否有遗漏的参数(比如Sort、Order等)。 - 启用Scrapy日志(
LOG_LEVEL=DEBUG),查看请求的状态码、响应内容,确认接口是否返回预期数据。 - 检查是否有反爬机制(比如Cookie验证),如果需要,可以在请求中携带从start_url获取的Cookie。
内容的提问来源于stack exchange,提问作者Amir Hamidi
相关产品推荐
相关产品推荐

