You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取Indiamart.com返回None,昨日正常今日失效求助

问题排查与解决方法

1. 爬虫入口方法错误

你的代码定义了request_header方法,但Scrapy爬虫默认只会自动调用start_requests方法,自定义的request_header不会被触发。这导致请求未携带自定义User-Agent,网站可能通过UA识别出爬虫,返回了状态码200但内容被隐藏或结构异常。

修复方法:
将request_header重命名为start_requests,同时注意start_urls是列表,需遍历处理每个链接(scrapy.Request的url参数需要字符串,不能直接传入列表):

def start_requests(self):
    for url in self.start_urls:
        yield scrapy.Request(url=url, callback=self.parse, headers={'User-Agent': self.user_agent})

2. XPath选择器失效

网站大概率夜间更新了HTML结构或类名,导致原XPath无法匹配目标元素。

验证方法:
使用Scrapy Shell测试当前页面结构:

scrapy shell "https://dir.indiamart.com/search.mp?ss=laptop&prdsrc=1&res=RC4" -s USER_AGENT="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/108.0.0.0 Safari/537.36"

进入Shell后执行原XPath语句,确认是否返回结果:

response.xpath("//span[@class='elps elps2 p10b0 fs14 tac mListNme']/a/text()").get()

若返回None,需打开浏览器开发者工具重新定位元素,更新XPath。

替代方案:
采用更宽松的选择器,先定位商品容器再提取内容:

products = response.xpath("//div[contains(@class, 'mList')]")
for product in products:
    title = product.xpath(".//a[contains(@class, 'mListNme')]/text()").get()
    related_link = product.xpath(".//a[contains(@class, 'mListNme')]/@href").get()
    yield {
        'titling': title,
        'rel_link': related_link
    }

3. 反爬机制触发

即使页面状态码为200,网站也可能通过以下手段返回空数据:

  • UA检测:确保请求携带真实浏览器UA
  • IP封禁:频繁请求可能被临时封禁,可更换IP或添加请求延迟
  • Cookie验证:部分网站需携带Cookie返回正常内容,可从浏览器开发者工具复制Cookie添加到请求头

添加请求延迟:
在settings.py中设置:

DOWNLOAD_DELAY = 2

完整修复代码示例

class IndiaSpider(scrapy.Spider):
    name = 'india'
    allowed_domains = ['indiamart.com']
    start_urls = ['https://dir.indiamart.com/search.mp?ss=laptop&prdsrc=1&res=RC4']

    user_agent = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/108.0.0.0 Safari/537.36'

    def start_requests(self):
        for url in self.start_urls:
            yield scrapy.Request(url=url, callback=self.parse, headers={'User-Agent': self.user_agent})

    def parse(self, response):
        products = response.xpath("//div[contains(@class, 'mList')]")
        for product in products:
            title = product.xpath(".//a[contains(@class, 'mListNme')]/text()").get(default='N/A')
            related_link = product.xpath(".//a[contains(@class, 'mListNme')]/@href").get(default='N/A')
            yield {
                'titling': title.strip() if title else title,
                'rel_link': response.urljoin(related_link) if related_link else related_link
            }

内容的提问来源于stack exchange,提问作者Sarfraz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 12:35:25