Python Requests POST请求失败,爬取Sayurbox蔬菜商品求助
Sayurbox蔬菜分类页面商品爬取问题修复
问题概述
尝试爬取Sayurbox蔬菜分类页面(https://www.sayurbox.com/category/vegetables-1-a0d03d59/sub-category/all)的所有商品,已获取页面源码中的authoId和DCId并填入代码,但请求始终无法返回200状态码,怀疑payload格式存在问题,尝试仅传入第一个getProducts请求但不确定是否可行。
核心问题分析
- 批量GraphQL请求限制:当前代码同时发送3个GraphQL请求(1个购物车计数+2个商品请求),部分服务器会限制批量请求的数量或格式,导致非200响应。
- 双重序列化错误:代码中先对payload执行
json.dumps(),再传给requests.post的json参数,会导致JSON被双重序列化,格式完全错误。 - 冗余请求干扰:
getCartItemCount请求并非爬取商品的必需请求,多余的请求可能触发服务器的风控机制。
修复方案
- 仅发送单个
getProducts请求:去掉购物车计数和硬编码分页的请求,只保留基础的商品查询请求。 - 修正序列化方式:直接将Python字典作为
json参数传入requests.post,无需手动执行json.dumps(),requests会自动完成正确的序列化。 - 补充完整请求头:除
authorization外,添加浏览器请求中常见的User-Agent、Referer等头,模拟真实浏览器请求,降低被拦截概率。 - 实现动态分页逻辑:根据响应中的
pageInfo.hasNextPage和pageInfo.endCursor,循环请求所有分页数据。
完整修复代码示例
import requests import pandas as pd from datetime import date today = str(date.today()) slugcat = "vegetables-1-a0d03d59" url = "https://www.sayurbox.com/graphql/v1?deduplicate=1" # 替换为你获取到的真实值 authoId = "YOUR_AUTHORIZATION_TOKEN" DCId = "YOUR_DELIVERY_CONFIG_ID" headers = { 'authorization': authoId, 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Referer': 'https://www.sayurbox.com/category/vegetables-1-a0d03d59/sub-category/all' } all_products = [] after_cursor = None while True: # 构建单个GraphQL请求payload payload = { "operationName": "getProducts", "variables": { "deliveryConfigId": DCId, "sortBy": "related_product", "isInstantDelivery": False, "slug": slugcat, "first": 12, "abTestFeatures": [], "after": after_cursor }, "query": """query getProducts($deliveryConfigId: ID!, $sortBy: CatalogueSortType!, $slug: String!, $after: String, $first: Int, $isInstantDelivery: Boolean, $abTestFeatures: [String!]) { productsByCategoryOrSubcategoryAndDeliveryConfig( deliveryConfigId: $deliveryConfigId sortBy: $sortBy slug: $slug after: $after first: $first isInstantDelivery: $isInstantDelivery abTestFeatures: $abTestFeatures ) { edges { node { ...ProductInfoFragment __typename } __typename } pageInfo { hasNextPage endCursor __typename } __typename } } fragment ProductInfoFragment on Product { id uuid displayName priceMin priceMax actualPriceMin actualPriceMax slug isStockAvailable quantitySoldFormatted productVariants { productVariant { skuCode variantName stockAvailable __typename } __typename } __typename }""" } response = requests.post(url, headers=headers, json=payload) # 检查响应状态 if response.status_code != 200: print(f"请求失败,状态码: {response.status_code}") print(response.text) break data = response.json() product_edges = data['data']['productsByCategoryOrSubcategoryAndDeliveryConfig']['edges'] all_products.extend([edge['node'] for edge in product_edges]) # 更新分页游标,判断是否继续请求 page_info = data['data']['productsByCategoryOrSubcategoryAndDeliveryConfig']['pageInfo'] if not page_info['hasNextPage']: break after_cursor = page_info['endCursor'] # 转换为DataFrame保存 df = pd.DataFrame(all_products) df.to_csv(f"sayurbox_vegetables_{today}.csv", index=False) print(f"爬取完成,共获取{len(all_products)}个商品")
注意事项
- 确保
authoId和DCId与浏览器请求中的值完全一致,注意不要遗漏前缀(如Bearer) - 控制请求频率,避免短时间内发送大量请求导致IP被封禁
- 若仍出现非200响应,可在浏览器的开发者工具中复制完整的请求头(包括
Cookie等)添加到代码的headers中
内容的提问来源于stack exchange,提问作者Hal
相关产品推荐
相关产品推荐

