使用Python recipe_scraper生成DataFrame时食材重复写入问题求助
问题解决:避免食谱数据重复写入DataFrame
问题根源
你的代码中**df_list.append(df_temp)被放在了内层的食材循环里**,每处理一种食材就把当前的df_temp(已经累积了之前的食材数据)添加到列表中。比如一个有13种食材的食谱,会被重复添加13次,最终合并后的DataFrame自然会出现大量重复数据。
修正方案
把df_list.append(df_temp)移到内层循环的外面,也就是处理完当前食谱的所有食材之后,再将完整的食谱DataFrame添加到列表中。同时调整DataFrame结构,让每个食谱对应一列存储食材用量,更贴合需求:
import pandas as pd from recipe_scrapers import scrape_me def replace_measurement_symbols(ingredients): """ 将分数符号转换为可计算的小数 参数: * ingredients: 食材列表对象 """ ingredients = [i.replace('¼', '0.25') for i in ingredients] ingredients = [i.replace('½', '0.5') for i in ingredients] ingredients = [i.replace('¾', '0.75') for i in ingredients] return ingredients def create_df(recipes): """ 生成包含所有食谱及其食材用量的DataFrame 参数: * recipes: 用户提供的食谱URL列表 """ df_list = [] for recipe in recipes: scraper = scrape_me(recipe) recipe_details = replace_measurement_symbols(scraper.ingredients()) # 提取食谱名称 recipe_name = recipe.split("https://www.hellofresh.nl/recipes/", 1)[1] recipe_name = recipe_name.rsplit('-', 1)[0] print(recipe_name) # 初始化当前食谱的临时字典,存储食材和对应的用量 recipe_data = {'Ingredients': [], recipe_name: []} for ingredient in recipe_details: try: ing_1 = ingredient.split("2 * ", 1)[1] ing_1 = ing_1.split(" ", 2) item = ing_1[2] measurement = ing_1[1] quantity = float(ing_1[0]) * 2 # 合并数量和单位,也可按需分开存储 full_measurement = f"{quantity} {measurement}" recipe_data['Ingredients'].append(item) recipe_data[recipe_name].append(full_measurement) except (ValueError, IndexError): # 捕获分割错误,避免程序中断 pass # 将当前食谱的字典转为DataFrame后添加到列表 df_temp = pd.DataFrame(recipe_data) df_list.append(df_temp) # 合并所有食谱DataFrame,按食材名称对齐去重 df = pd.concat(df_list, axis=0).groupby('Ingredients').first().reset_index() return df def main(): recipes = [ 'https://www.hellofresh.nl/recipes/luxe-burger-met-truffeltapenade-en-portobello-63ad875558b39f3da6083acd', 'https://www.hellofresh.nl/recipes/chicken-parmigiana-623c51bd7ed5c074f51bbb10', 'https://www.hellofresh.nl/recipes/quiche-met-broccoli-en-oude-kaas-628665b01dea7b8f5009b248', ] df = create_df(recipes) print(df) if __name__ == "__main__": main()
关键调整点
- 移动append位置:将
df_list.append(df_temp)从内层食材循环移到外层食谱循环末尾,确保每个食谱只被添加一次。 - 优化数据收集:用字典先统一收集当前食谱的食材和用量,再转为DataFrame,避免逐行添加的低效操作。
- 对齐去重:通过
groupby('Ingredients').first()确保同一食材在不同食谱中的数据正确对齐,消除重复行。 - 增强异常处理:添加
IndexError捕获,避免食材字符串分割失败导致程序崩溃。
内容的提问来源于stack exchange,提问作者June Smith
相关产品推荐
相关产品推荐

