Vivino数据爬取返回记录数远低于预期的技术问题
解决Vivino API爬取记录数不符的问题
你遇到的问题核心是Vivino的/api/explore/explore接口默认带隐性过滤规则,再加上分页参数设置不合理,导致返回记录远少于实际的wines_count。以下是具体解决思路和优化后的代码:
1. 问题根源拆解
- 接口默认仅返回有公开价格且有一定评分数量的酒款,直接过滤了大量无价格、评分极少的酒款
- 默认单页仅返回25条数据,你固定循环到99页,最多只能拿到2475条,远达不到匹配总数
- 频繁无延迟请求可能触发反爬限制,导致接口返回数据不全
2. 具体解决措施
- 用
per_page参数提高单页返回量(建议设为100,过大易触发限制) - 添加
filters=all和min_ratings_count=0取消默认过滤规则 - 根据接口返回的
records_matched动态计算总页数,避免无效请求 - 增加随机请求延迟,规避反爬限制
- 处理字段为空的异常情况,避免代码报错中断
优化后的代码示例
import json import pandas as pd import requests import time import random # 初始化存储DataFrame wine_cols = ['Winery','Wine','Rating','Review_Count','Region','SubRegion', 'Price', 'id'] wine_df = pd.DataFrame(columns=wine_cols) # 请求基础配置 base_url = "https://www.vivino.com/api/explore/explore" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:117.0) Gecko/20100101 Firefox/117.0" } # 先请求第一页获取总匹配数 first_req = requests.get( base_url, params={ "country_codes[]": "it", "per_page": 100, "filters": "all", "min_ratings_count": 0 }, headers=headers ) first_data = first_req.json() total_matched = first_data["explore_vintage"]["records_matched"] total_pages = (total_matched + 99) // 100 # 向上取整计算总页数 # 循环爬取所有页面 for page in range(1, total_pages + 1): print(f"爬取第 {page}/{total_pages} 页") # 随机延迟1-3秒,避免反爬 time.sleep(random.uniform(1, 3)) req = requests.get( base_url, params={ "country_codes[]": "it", "page": page, "per_page": 100, "filters": "all", "min_ratings_count": 0 }, headers=headers ) # 处理空数据情况 matches = req.json().get("explore_vintage", {}).get("matches") if not matches: print(f"第 {page} 页无有效数据,跳过") continue # 提取数据并处理空字段 results = [ ( t["vintage"]["wine"]["winery"]["name"], f'{t["vintage"]["wine"]["name"]} {t["vintage"]["year"] if t["vintage"]["year"] else "无年份"}', t["vintage"]["statistics"].get("ratings_average", 0), t["vintage"]["statistics"].get("ratings_count", 0), t["vintage"]["wine"]["region"]["country"]["name"], t["vintage"]["wine"]["region"]["name"], t["price"]["amount"] if t.get("price") else 0, t["vintage"]["wine"]["id"] ) for t in matches ] temp_df = pd.DataFrame(results, columns=wine_cols) wine_df = pd.concat([wine_df, temp_df], ignore_index=True) # 用concat替代已弃用的append print(f"爬取完成,共获取 {len(wine_df)} 条数据") wine_df.to_csv("italian_wines.csv", index=False)
补充说明
- 即使调整参数,也可能无法获取全部38万条数据,因为Vivino会隐藏未上架、已下架的酒款,API本身也有返回上限
- 如果出现403/429错误,说明触发反爬限制,需延长延迟时间或使用代理IP
- 可尝试更换
sort_by参数(如most_popular、price),不同排序可能覆盖更多酒款池
内容的提问来源于stack exchange,提问作者TheDalek
相关产品推荐
相关产品推荐

