You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy爬取无限滚动页面动态商品数据

Trendyol无限滚动分页商品爬取实现方案

直接调用站点公开的无限滚动接口爬取,不需要模拟浏览器滚动操作,爬取效率和稳定性远高于Selenium等自动化方案,具体实现如下:

接口规律拆解

  • 接口基础路径为https://public.trendyol.com/discovery-web-searchgw-service/v2/api/infinite-scroll/,后缀直接拼接分类页URL末尾的分类标识段即可,例如洗衣液分类页路径为camasir-deterjani-x-c108713,对应接口路径就是基础路径加该段标识。
  • 分页核心参数为pi,对应页码值,首屏加载内容为pi=1的返回结果,每向下滚动加载一页,参数值加1。
  • 其余参数(地区、用户标识、算法ID等)均为固定值,直接复用抓包获取的参数即可,短时间内不会变更。
  • 接口校验请求头中的Referer字段,值必须为对应分类的前台页面地址,否则会直接返回403拦截。

代码改造示例

直接替换你原有爬虫的对应逻辑即可,原有Item字段定义可以完全保留:

import scrapy
import json
import datetime
from your_project.items import TeknosaItem

class TrendyolSpider(scrapy.Spider):
    name = 'trendyol'
    # 开自动限速,避免请求过快被封
    custom_settings = {
        'DOWNLOAD_DELAY': 2,
        'RETRY_TIMES': 3
    }

    def start_requests(self):
        # 保留原有分类列表
        category_list = [
            'camasir-deterjani-x-c108713',
            'yumusaticilar-x-c103814',
            'camasir-suyu-x-c103812',
            'camasir-leke-cikaricilar-x-c103810',
            'camasir-yan-urun-x-c105534',
            'kirec-onleyici-x-c103806',
            'makine-kirec-onleyici-ve-temizleyici-x-c144512'
        ]
        # 接口固定参数,直接复用抓包结果
        base_params = {
            "culture": "tr-TR",
            "userGenderId": "1",
            "pId": "0",
            "scoringAlgorithmId": "2",
            "categoryRelevancyEnabled": "false",
            "isLegalRequirementConfirmed": "false",
            "searchStrategyType": "DEFAULT",
            "productStampType": "TypeA",
            "fixSlotProductAdsIncluded": "false"
        }
        api_domain = "https://public.trendyol.com/discovery-web-searchgw-service/v2/api/infinite-scroll/"
        page_domain = "https://www.trendyol.com/"

        for category in category_list:
            # 初始从第一页开始爬,后续碰到空返回自动终止
            current_pi = 1
            params = base_params.copy()
            params["pi"] = str(current_pi)
            headers = {
                "Referer": f"{page_domain}{category}",
                "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36"
            }
            yield scrapy.FormRequest(
                url=f"{api_domain}{category}",
                method="GET",
                formdata=params,
                headers=headers,
                callback=self.parse_product,
                cb_kwargs={
                    "category": category,
                    "base_params": base_params,
                    "api_domain": api_domain,
                    "page_domain": page_domain,
                    "current_pi": current_pi
                }
            )

    def parse_product(self, response, category, base_params, api_domain, page_domain, current_pi):
        resp_data = json.loads(response.text)
        product_list = resp_data.get("result", {}).get("products", [])
        # 当前页无商品,说明该分类已爬完,终止翻页
        if not product_list:
            return
        
        # 解析当前页商品,复用原有字段逻辑
        for p in product_list:
            item = TeknosaItem()
            item['rowid'] = hash(str(datetime.datetime.now()) + str(p["id"]))
            item['date'] = str(datetime.datetime.now())
            item['listing_id'] = p["id"]
            item['product_id'] = p["id"]
            item['product_name'] = p["name"]
            item['price'] = p["price"]["sellingPrice"]
            # 接口返回的url是相对路径,拼接主域名
            item['url'] = f"{page_domain}{p['url'].lstrip('/')}"
            yield item
        
        # 构造下一页请求
        next_pi = current_pi + 1
        next_params = base_params.copy()
        next_params["pi"] = str(next_pi)
        headers = {
            "Referer": f"{page_domain}{category}",
            "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36"
        }
        yield scrapy.FormRequest(
            url=f"{api_domain}{category}",
            method="GET",
            formdata=next_params,
            headers=headers,
            callback=self.parse_product,
            cb_kwargs={
                "category": category,
                "base_params": base_params,
                "api_domain": api_domain,
                "page_domain": page_domain,
                "current_pi": next_pi
            }
        )

注意事项

  • 不需要用BeautifulSoup或者正则匹配页面源码提取数据,接口直接返回结构化JSON,解析稳定性更高,不会因为前端页面结构改版轻易失效。
  • 不要删除Referer请求头,否则会被站点反爬拦截;如果出现403/429错误,可以降低请求频率、轮换User-Agent、搭配代理池使用。
  • 代码里没有设置固定的爬取页数上限,碰到空商品列表会自动终止当前分类的爬取,不会产生无效请求。
  • 如果后续接口返回参数错误,打开浏览器开发者工具抓一次最新的无限滚动请求,更新base_params里的固定参数即可。

内容的提问来源于stack exchange,提问作者fatma hilal erol

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 17:30:47