You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过POST请求获取Aldi网站的Pagination.PagesCount分页数据

爬取Aldi爱尔兰生鲜网站时获取最大页码的问题

我正在爬取Aldi爱尔兰生鲜网站(groceries.aldi.ie)的商品信息,需要获取分类页面的最大页码来控制爬虫停止时机。近期网站更新了页码展示逻辑,现在必须通过POST请求获取最大页码,但旧代码依赖的Pagination.PagesCount数据找不到对应的接口。

以下是之前用来获取页面HTML的代码(无法提取页码):

from scrapy import Selector
import requests

url = 'https://groceries.aldi.ie/en-GB/chilled-food/cheese?origin=dropdown&c1=shopgroceries&c2=chilled-food&c3=cheese&clickedon=cheese'

headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:84.0) Gecko/20100101 Firefox/84.0',
           'Accept-Language': 'en-GB,en;q=0.5',
           'Referer': 'https://google.com',
           'DNT': '1'}

html = requests.get(url, headers=headers).content
sel = Selector(text=html)

print(sel)

我之前成功通过POST请求获取过商品定价数据(代码如下),但检查页面源码和XHR请求后,始终找不到返回Pagination.PagesCount的接口,求解决方法:

import json

headers = {
    "authority": "groceries.aldi.ie",
    "pragma": "no-cache",
    "cache-control": "no-cache",
    "sec-ch-ua": "\" Not;A Brand\";v=\"99\", \"Google Chrome\";v=\"91\", \"Chromium\";v=\"91\"",
    "accept-language": "en-GB",
    "sec-ch-ua-mobile": "?0",
    "user-agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.164 Safari/537.36",
    "websiteid": "a763fb4a-0224-4ca8-bdaa-a33a4b47a026",
    "content-type": "application/json",
    "accept": "application/json, text/javascript, */*; q=0.01",
    "x-requested-with": "XMLHttpRequest",
    "origin": "https://groceries.aldi.ie",
    "sec-fetch-site": "same-origin",
    "sec-fetch-mode": "cors",
    "sec-fetch-dest": "empty",
    "referer": "https://groceries.aldi.ie/en-GB/chilled-food/cheese?origin=dropdown&c1=shopgroceries&c2=chilled-food&c3=cheese&clickedon=cheese"
}

body = '{"products":["4088600284026","5391528370382","5391528372836","5391528372850","5391528372874","4088600298696","4088600103709","4088600388700","5035766046028","5000213021934","4088600012551","4088600325934","4088600300153","25389111","4072700001171","4088600012537","4088600012544","4088600013138","4088600013145","4088600103525","4088600103532","4088600103570","4088600103600","4088600135182","4088600141848","4088600142050","4088600158105","4088600217024","4088600217208","4088600241302","4088600249292","4088600249308","4088600280615","4088600281445","4088600283043","4088600284088","4088600295688","4088600295800","4088600295817","4088600303925"]}'

url = 'https://groceries.aldi.ie/api/product/calculatePrices'
response = requests.post(url=url, headers=headers,data=body)

data = response.text
data = data.replace("'", '\"')
d = json.loads(data)

print(d)

解决方案

核心思路

该网站的分页和商品列表数据来自https://groceries.aldi.ie/api/products/search接口,通过POST请求传递分类、分页参数即可获取包含总页数的响应数据。

步骤与代码示例

  1. 构造请求参数:请求体需包含分类ID、每页商品数、页码等关键信息,分类ID可从页面URL路径提取(如奶酪分类为chilled-food/cheese)。
  2. 发送POST请求:使用正确的请求头(可复用之前定价接口的大部分头信息),发送JSON格式的请求体。
  3. 提取最大页码:响应JSON中pagination.totalPages字段即为所需的最大页码。
import requests
import json

headers = {
    "authority": "groceries.aldi.ie",
    "pragma": "no-cache",
    "cache-control": "no-cache",
    "sec-ch-ua": "\"Not/A)Brand\";v=\"99\", \"Google Chrome\";v=\"115\", \"Chromium\";v=\"115\"",
    "accept-language": "en-GB",
    "sec-ch-ua-mobile": "?0",
    "user-agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/115.0.0.0 Safari/537.36",
    "websiteid": "a763fb4a-0224-4ca8-bdaa-a33a4b47a026",
    "content-type": "application/json",
    "accept": "application/json, text/javascript, */*; q=0.01",
    "x-requested-with": "XMLHttpRequest",
    "origin": "https://groceries.aldi.ie",
    "sec-fetch-site": "same-origin",
    "sec-fetch-mode": "cors",
    "sec-fetch-dest": "empty",
    "referer": "https://groceries.aldi.ie/en-GB/chilled-food/cheese?origin=dropdown&c1=shopgroceries&c2=chilled-food&c3=cheese&clickedon=cheese"
}

# 构造请求体,以奶酪分类为例
payload = {
    "categoryIds": ["chilled-food/cheese"],
    "pageSize": 24,
    "pageNumber": 1,
    "sortBy": "Popularity",
    "sortDirection": "Descending",
    "filters": [],
    "searchTerm": ""
}

url = "https://groceries.aldi.ie/api/products/search"
response = requests.post(url, headers=headers, json=payload)
data = response.json()

# 提取最大页码
max_pages = data["pagination"]["totalPages"]
print(f"最大页码:{max_pages}")

注意事项

  • 若websiteid过期,可从浏览器开发者工具的请求头中抓取最新值。
  • 该接口同时返回当前页的商品数据,可直接用于商品信息爬取,无需再解析HTML页面。
  • 不同分类的categoryIds需对应替换,可从页面URL或源码中提取。

内容的提问来源于stack exchange,提问作者Oxin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 19:24:58