You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Vivino数据爬取返回记录数远低于预期的技术问题

解决Vivino API爬取记录数不符的问题

你遇到的问题核心是Vivino的/api/explore/explore接口默认带隐性过滤规则,再加上分页参数设置不合理,导致返回记录远少于实际的wines_count。以下是具体解决思路和优化后的代码:

1. 问题根源拆解

  • 接口默认仅返回有公开价格且有一定评分数量的酒款,直接过滤了大量无价格、评分极少的酒款
  • 默认单页仅返回25条数据,你固定循环到99页,最多只能拿到2475条,远达不到匹配总数
  • 频繁无延迟请求可能触发反爬限制,导致接口返回数据不全

2. 具体解决措施

  • 用per_page参数提高单页返回量(建议设为100,过大易触发限制)
  • 添加filters=all和min_ratings_count=0取消默认过滤规则
  • 根据接口返回的records_matched动态计算总页数,避免无效请求
  • 增加随机请求延迟,规避反爬限制
  • 处理字段为空的异常情况,避免代码报错中断

优化后的代码示例

import json
import pandas as pd
import requests
import time
import random

# 初始化存储DataFrame
wine_cols = ['Winery','Wine','Rating','Review_Count','Region','SubRegion', 'Price', 'id']
wine_df = pd.DataFrame(columns=wine_cols)

# 请求基础配置
base_url = "https://www.vivino.com/api/explore/explore"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:117.0) Gecko/20100101 Firefox/117.0"
}

# 先请求第一页获取总匹配数
first_req = requests.get(
    base_url,
    params={
        "country_codes[]": "it",
        "per_page": 100,
        "filters": "all",
        "min_ratings_count": 0
    },
    headers=headers
)
first_data = first_req.json()
total_matched = first_data["explore_vintage"]["records_matched"]
total_pages = (total_matched + 99) // 100  # 向上取整计算总页数

# 循环爬取所有页面
for page in range(1, total_pages + 1):
    print(f"爬取第 {page}/{total_pages} 页")
    # 随机延迟1-3秒,避免反爬
    time.sleep(random.uniform(1, 3))
    
    req = requests.get(
        base_url,
        params={
            "country_codes[]": "it",
            "page": page,
            "per_page": 100,
            "filters": "all",
            "min_ratings_count": 0
        },
        headers=headers
    )
    
    # 处理空数据情况
    matches = req.json().get("explore_vintage", {}).get("matches")
    if not matches:
        print(f"第 {page} 页无有效数据,跳过")
        continue
    
    # 提取数据并处理空字段
    results = [
        (
            t["vintage"]["wine"]["winery"]["name"],
            f'{t["vintage"]["wine"]["name"]} {t["vintage"]["year"] if t["vintage"]["year"] else "无年份"}',
            t["vintage"]["statistics"].get("ratings_average", 0),
            t["vintage"]["statistics"].get("ratings_count", 0),
            t["vintage"]["wine"]["region"]["country"]["name"],
            t["vintage"]["wine"]["region"]["name"],
            t["price"]["amount"] if t.get("price") else 0,
            t["vintage"]["wine"]["id"]
        )
        for t in matches
    ]
    
    temp_df = pd.DataFrame(results, columns=wine_cols)
    wine_df = pd.concat([wine_df, temp_df], ignore_index=True)  # 用concat替代已弃用的append

print(f"爬取完成,共获取 {len(wine_df)} 条数据")
wine_df.to_csv("italian_wines.csv", index=False)

补充说明

  • 即使调整参数,也可能无法获取全部38万条数据,因为Vivino会隐藏未上架、已下架的酒款,API本身也有返回上限
  • 如果出现403/429错误,说明触发反爬限制,需延长延迟时间或使用代理IP
  • 可尝试更换sort_by参数(如most_popular、price),不同排序可能覆盖更多酒款池

内容的提问来源于stack exchange,提问作者TheDalek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 20:03:13