You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy提交Mercado Libre搜索表单时遇502错误求助

问题分析与解决方法

核心问题

你遇到的502错误主要源于两个关键问题:

  1. 表单参数错误:你使用的cb1-edit是搜索输入框的id属性,而非表单提交时实际需要的name属性(正确参数为q),导致生成了格式错误的URL(末尾的%3E是非法字符转义)。
  2. 反爬机制拦截:Scrapy默认的User-Agent容易被识别为爬虫,触发服务器返回错误。

修正后的代码方案

方案1:直接构造搜索URL(更简洁)

跳过表单提交环节,直接生成合法的搜索请求:

class MlSpider(scrapy.Spider):
    name = 'ml'
    allowed_domains = ['mercadolivre.com.br']
    search_keyword = 'smartphone'

    def start_requests(self):
        # 直接构造搜索URL,使用正确的参数q
        search_url = f'https://www.mercadolivre.com.br/search?q={self.search_keyword}'
        yield scrapy.Request(
            url=search_url,
            callback=self.scrape_data,
            # 添加真实浏览器的User-Agent,避免被反爬拦截
            headers={
                'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
            }
        )

    def scrape_data(self, response):
        for element in response.xpath('//li[@class="ui-search-layout__item shops__layout-item"]'):
            # 使用相对路径(.//)确保仅获取当前item下的内容
            item_name = element.xpath('.//h2[@class="ui-search-item__title shops__item-title"]/text()').get()
            price = element.xpath('.//span[@class="andes-money-amount__fraction"]/text()').get()
            link = element.xpath('./a/@href').get()

            yield {
                "item": item_name,
                "price": price,
                "link": link
            }

方案2:修复FormRequest提交

如果坚持使用表单提交,需修正参数和表单定位:

class MlSpider(scrapy.Spider):
    name = 'ml'
    allowed_domains = ['mercadolivre.com.br']
    start_urls = ['https://www.mercadolivre.com.br/']

    def parse(self, response):
        return scrapy.FormRequest.from_response(
            response,
            # 精准定位搜索表单
            formxpath='//form[@role="search"]',
            # 使用正确的表单参数q
            formdata={'q': 'smartphone'},
            callback=self.scrape_data,
            headers={
                'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
            }
        )

    def scrape_data(self, response):
        for element in response.xpath('//li[@class="ui-search-layout__item shops__layout-item"]'):
            item_name = element.xpath('.//h2[@class="ui-search-item__title shops__item-title"]/text()').get()
            price = element.xpath('.//span[@class="andes-money-amount__fraction"]/text()').get()
            link = element.xpath('./a/@href').get()

            yield {
                "item": item_name,
                "price": price,
                "link": link
            }

关键修正点说明

  • 参数修正:将cb1-edit替换为搜索表单实际使用的q参数,确保URL格式合法。
  • User-Agent设置:添加真实浏览器的User-Agent,绕过基础反爬检测。
  • XPath优化:使用相对路径(.//)替代绝对路径(//),避免循环中重复获取同一个元素的内容。
  • 价格提取优化:直接定位价格文本节点,而非获取整个HTML块,返回更干净的数据。

内容的提问来源于stack exchange,提问作者Kabllez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 11:20:28