You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用爬虫获取的rating_value替换CSV中的rating_score字段值

问题:爬虫获取的评分无法替换CSV中原有评分字段

现有Python脚本可读取venues.csv与beer_full.csv,通过爬虫从venues.csv的链接获取啤酒ID,匹配beer_full.csv的数据后写入menu_beers.csv。目前已能成功爬取到rating_value,但无法用它替换beer_full.csv里的rating_score字段,尝试row['rating_score'] = rating_value无效。

核心问题分析

原来的逻辑只收集了啤酒ID列表,没有保存每个ID对应的爬取评分,后续直接从all_beer中筛选数据时,没有将爬取到的评分与对应啤酒关联,所以替换操作无效。

修改方案

  1. 将存储啤酒ID的全局变量从列表改为字典,用来保存{beer_id: rating_value}的映射关系
  2. 修改get_menu_beers函数,把爬取到的啤酒ID和对应评分存入字典
  3. 在筛选出当前酒吧的啤酒数据后,利用字典映射替换rating_score字段的值

修改后的完整代码

import pandas as pd
import warnings
import requests
from bs4 import BeautifulSoup
import re
import os
import time
import sys


def resource_path(relative_path: str) -> str:
    try:
        base_path = sys._MEIPASS
    except Exception:
        base_path = os.path.dirname(__file__)
    return os.path.join(base_path, relative_path)


warnings.simplefilter(action="ignore", category=FutureWarning)


df = pd.read_csv(resource_path("venues.csv"))
all_beer = pd.read_csv(resource_path("beer_full.csv"), encoding="ISO-8859-1")
fname = "menu_beers.csv"


def get_menu_beers(soup):
    global bar_beer_rating_map
    beers_all = soup.find_all("ul", {"class": "menu-section-list"})
    for beer_group in beers_all:
        beers = beer_group.find_all("li")
        for beer in beers:
            details = beer.find("div", {"class": "beer-details"})
            a_href = details.find("a", {"class": "track-click"}).get("href")
            id_num = re.findall(r"\d+", a_href)
            beer_id = int(id_num[-1])
            # 存储啤酒ID和对应爬取的评分
            rating_value = details.find('div', {'class': 'caps small'})['data-rating']
            bar_beer_rating_map[beer_id] = float(rating_value)


for index, row in df.iterrows():
    url_base = (
        "https://example.com/v/" + str(row["venue_slug"]) + "/" + str(row["venue_id"])
    )
    url = url_base + "/beers"
    # 初始化字典,替代原来的列表
    bar_beer_rating_map = {}
    response = requests.get(url, headers={"User-agent": "Mozilla/5.0"})
    if response.status_code == 200:
        soup = BeautifulSoup(response.content, "html.parser")
        try:
            try:
                select_options = soup.find_all("select", {"class": "menu-selector"})
                if len(select_options) > 0:
                    options_list = select_options[0].find_all("option")
                    menu_ids = []
                    for option in options_list:
                        menu_ids.append(int(option["value"]))
                    menu_urls = []
                    for menu_id in menu_ids:
                        menu_url = str(url_base) + "?menu_id=" + str(menu_id)
                        menu_urls.append(menu_url)
                    for url in menu_urls:
                        res = requests.get(url, headers={"User-agent": "Mozilla/5.0"})
                        s = BeautifulSoup(res.text, "html.parser")
                        get_menu_beers(s)
                else:
                    get_menu_beers(soup)
            except:
                print("Error at " + str(row["venue_name"]))

            # 获取当前酒吧的所有啤酒ID
            bar_beer_ids = list(bar_beer_rating_map.keys())
            location_distance = row["distance_from_point"]

            bar_beers = all_beer.loc[all_beer["beer_id"].isin(bar_beer_ids)].copy()
            # 替换rating_score字段
            bar_beers['rating_score'] = bar_beers['beer_id'].map(bar_beer_rating_map)
            
            bar_beers.insert(0, "venue_name", str(row["venue_name"]))
            bar_beers.insert(1, "location_distance", round(location_distance, 2))
            del bar_beers["beer_id"]

            bar_beers.to_csv(fname, header=False, index=False, mode="a")
            print("Fetching menu for " + str(row["venue_name"]))
        except:
            print("No menu for " + str(row["venue_name"]))
            pass

    elif response.status_code == 429:
        print("Server currently receiving too many requests. Pausing for 1 minute")
        time.sleep(60)
        print("Resuming...")
        continue
    else:
        print("Could not process " + str(row["venue_name"]))


print("\n\n\n\n\n---------- DONE ------------")
print("Saving results to csv...")
print(
    "Column headers for csv in order are: venue_name,location_distance,beer_name,brewery_name,brewery_id,type_name,beer_abv,rating_score,rating_count,in_production"
)

关键修改点说明

  • 将bar_beer_ids列表改为bar_beer_rating_map字典,存储每个啤酒ID对应的爬取评分
  • 在get_menu_beers中,不再向列表追加ID,而是将ID和评分存入字典
  • 筛选出bar_beers后,使用map方法将字典中的评分映射到rating_score字段,调用copy()避免SettingWithCopyWarning
  • 修正了输出的列头说明,确保和实际输出一致

内容的提问来源于stack exchange,提问作者Tendekai Muchenje

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 04:55:41