You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python recipe_scraper生成DataFrame时食材重复写入问题求助

问题解决:避免食谱数据重复写入DataFrame

问题根源

你的代码中**df_list.append(df_temp)被放在了内层的食材循环里**,每处理一种食材就把当前的df_temp(已经累积了之前的食材数据)添加到列表中。比如一个有13种食材的食谱,会被重复添加13次,最终合并后的DataFrame自然会出现大量重复数据。

修正方案

把df_list.append(df_temp)移到内层循环的外面,也就是处理完当前食谱的所有食材之后,再将完整的食谱DataFrame添加到列表中。同时调整DataFrame结构,让每个食谱对应一列存储食材用量,更贴合需求:

import pandas as pd
from recipe_scrapers import scrape_me


def replace_measurement_symbols(ingredients):
    """
    将分数符号转换为可计算的小数
    参数:
    * ingredients: 食材列表对象
    """
    ingredients = [i.replace('¼', '0.25') for i in ingredients]
    ingredients = [i.replace('½', '0.5') for i in ingredients]
    ingredients = [i.replace('¾', '0.75') for i in ingredients]

    return ingredients


def create_df(recipes):
    """
    生成包含所有食谱及其食材用量的DataFrame
    参数:
    * recipes: 用户提供的食谱URL列表
    """
    df_list = []

    for recipe in recipes:
        scraper = scrape_me(recipe)
        recipe_details = replace_measurement_symbols(scraper.ingredients())

        # 提取食谱名称
        recipe_name = recipe.split("https://www.hellofresh.nl/recipes/", 1)[1]
        recipe_name = recipe_name.rsplit('-', 1)[0]
        print(recipe_name)

        # 初始化当前食谱的临时字典,存储食材和对应的用量
        recipe_data = {'Ingredients': [], recipe_name: []}

        for ingredient in recipe_details:
            try:
                ing_1 = ingredient.split("2 * ", 1)[1]
                ing_1 = ing_1.split(" ", 2)

                item = ing_1[2]
                measurement = ing_1[1]
                quantity = float(ing_1[0]) * 2
                # 合并数量和单位,也可按需分开存储
                full_measurement = f"{quantity} {measurement}"

                recipe_data['Ingredients'].append(item)
                recipe_data[recipe_name].append(full_measurement)
            except (ValueError, IndexError):
                # 捕获分割错误,避免程序中断
                pass

        # 将当前食谱的字典转为DataFrame后添加到列表
        df_temp = pd.DataFrame(recipe_data)
        df_list.append(df_temp)

    # 合并所有食谱DataFrame,按食材名称对齐去重
    df = pd.concat(df_list, axis=0).groupby('Ingredients').first().reset_index()

    return df


def main():
    recipes = [
        'https://www.hellofresh.nl/recipes/luxe-burger-met-truffeltapenade-en-portobello-63ad875558b39f3da6083acd',
        'https://www.hellofresh.nl/recipes/chicken-parmigiana-623c51bd7ed5c074f51bbb10',
        'https://www.hellofresh.nl/recipes/quiche-met-broccoli-en-oude-kaas-628665b01dea7b8f5009b248',
    ]

    df = create_df(recipes)
    print(df)


if __name__ == "__main__":
    main()

关键调整点

  1. 移动append位置:将df_list.append(df_temp)从内层食材循环移到外层食谱循环末尾,确保每个食谱只被添加一次。
  2. 优化数据收集:用字典先统一收集当前食谱的食材和用量,再转为DataFrame,避免逐行添加的低效操作。
  3. 对齐去重:通过groupby('Ingredients').first()确保同一食材在不同食谱中的数据正确对齐,消除重复行。
  4. 增强异常处理:添加IndexError捕获,避免食材字符串分割失败导致程序崩溃。

内容的提问来源于stack exchange,提问作者June Smith

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 01:40:33