如何通过POST请求获取Aldi网站的Pagination.PagesCount分页数据
爬取Aldi爱尔兰生鲜网站时获取最大页码的问题
我正在爬取Aldi爱尔兰生鲜网站(groceries.aldi.ie)的商品信息,需要获取分类页面的最大页码来控制爬虫停止时机。近期网站更新了页码展示逻辑,现在必须通过POST请求获取最大页码,但旧代码依赖的Pagination.PagesCount数据找不到对应的接口。
以下是之前用来获取页面HTML的代码(无法提取页码):
from scrapy import Selector import requests url = 'https://groceries.aldi.ie/en-GB/chilled-food/cheese?origin=dropdown&c1=shopgroceries&c2=chilled-food&c3=cheese&clickedon=cheese' headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:84.0) Gecko/20100101 Firefox/84.0', 'Accept-Language': 'en-GB,en;q=0.5', 'Referer': 'https://google.com', 'DNT': '1'} html = requests.get(url, headers=headers).content sel = Selector(text=html) print(sel)
我之前成功通过POST请求获取过商品定价数据(代码如下),但检查页面源码和XHR请求后,始终找不到返回Pagination.PagesCount的接口,求解决方法:
import json headers = { "authority": "groceries.aldi.ie", "pragma": "no-cache", "cache-control": "no-cache", "sec-ch-ua": "\" Not;A Brand\";v=\"99\", \"Google Chrome\";v=\"91\", \"Chromium\";v=\"91\"", "accept-language": "en-GB", "sec-ch-ua-mobile": "?0", "user-agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.164 Safari/537.36", "websiteid": "a763fb4a-0224-4ca8-bdaa-a33a4b47a026", "content-type": "application/json", "accept": "application/json, text/javascript, */*; q=0.01", "x-requested-with": "XMLHttpRequest", "origin": "https://groceries.aldi.ie", "sec-fetch-site": "same-origin", "sec-fetch-mode": "cors", "sec-fetch-dest": "empty", "referer": "https://groceries.aldi.ie/en-GB/chilled-food/cheese?origin=dropdown&c1=shopgroceries&c2=chilled-food&c3=cheese&clickedon=cheese" } body = '{"products":["4088600284026","5391528370382","5391528372836","5391528372850","5391528372874","4088600298696","4088600103709","4088600388700","5035766046028","5000213021934","4088600012551","4088600325934","4088600300153","25389111","4072700001171","4088600012537","4088600012544","4088600013138","4088600013145","4088600103525","4088600103532","4088600103570","4088600103600","4088600135182","4088600141848","4088600142050","4088600158105","4088600217024","4088600217208","4088600241302","4088600249292","4088600249308","4088600280615","4088600281445","4088600283043","4088600284088","4088600295688","4088600295800","4088600295817","4088600303925"]}' url = 'https://groceries.aldi.ie/api/product/calculatePrices' response = requests.post(url=url, headers=headers,data=body) data = response.text data = data.replace("'", '\"') d = json.loads(data) print(d)
解决方案
核心思路
该网站的分页和商品列表数据来自https://groceries.aldi.ie/api/products/search接口,通过POST请求传递分类、分页参数即可获取包含总页数的响应数据。
步骤与代码示例
- 构造请求参数:请求体需包含分类ID、每页商品数、页码等关键信息,分类ID可从页面URL路径提取(如奶酪分类为
chilled-food/cheese)。 - 发送POST请求:使用正确的请求头(可复用之前定价接口的大部分头信息),发送JSON格式的请求体。
- 提取最大页码:响应JSON中
pagination.totalPages字段即为所需的最大页码。
import requests import json headers = { "authority": "groceries.aldi.ie", "pragma": "no-cache", "cache-control": "no-cache", "sec-ch-ua": "\"Not/A)Brand\";v=\"99\", \"Google Chrome\";v=\"115\", \"Chromium\";v=\"115\"", "accept-language": "en-GB", "sec-ch-ua-mobile": "?0", "user-agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/115.0.0.0 Safari/537.36", "websiteid": "a763fb4a-0224-4ca8-bdaa-a33a4b47a026", "content-type": "application/json", "accept": "application/json, text/javascript, */*; q=0.01", "x-requested-with": "XMLHttpRequest", "origin": "https://groceries.aldi.ie", "sec-fetch-site": "same-origin", "sec-fetch-mode": "cors", "sec-fetch-dest": "empty", "referer": "https://groceries.aldi.ie/en-GB/chilled-food/cheese?origin=dropdown&c1=shopgroceries&c2=chilled-food&c3=cheese&clickedon=cheese" } # 构造请求体,以奶酪分类为例 payload = { "categoryIds": ["chilled-food/cheese"], "pageSize": 24, "pageNumber": 1, "sortBy": "Popularity", "sortDirection": "Descending", "filters": [], "searchTerm": "" } url = "https://groceries.aldi.ie/api/products/search" response = requests.post(url, headers=headers, json=payload) data = response.json() # 提取最大页码 max_pages = data["pagination"]["totalPages"] print(f"最大页码:{max_pages}")
注意事项
- 若
websiteid过期,可从浏览器开发者工具的请求头中抓取最新值。 - 该接口同时返回当前页的商品数据,可直接用于商品信息爬取,无需再解析HTML页面。
- 不同分类的
categoryIds需对应替换,可从页面URL或源码中提取。
内容的提问来源于stack exchange,提问作者Oxin
相关产品推荐
相关产品推荐

