使用BeautifulSoup爬取Allrecipes食谱:食材与步骤解析失败求助
问题
我正在编写程序从allrecipes.com的食谱页面提取标题、URL、食材和烹饪步骤,目前先针对单个页面测试,但食材和步骤的解析模块有问题——HTML中这两部分是列表格式,我无法将其解析为可用数据。作为BS4新手,附上代码及输出结果寻求帮助:
原代码:
import requests from bs4 import BeautifulSoup import json def scrape_recipe_data(url): # Make an HTTP request to the provided URL response = requests.get(url) # Check if the request was successful (status code 200) if response.status_code == 200: soup = BeautifulSoup(response.text, 'html.parser') # Extract title title_tag = soup.title title_text = title_tag.text if title_tag else "No title found in the HTML document" print("Title:", title_text) ingredients_element = soup.find('script', {'type': 'application/ld+json', 'title': 'recipeIngredient'}) if ingredients_element: ingredients_data = json.loads(ingredients_element.string) ingredients = ingredients_data if isinstance(ingredients_data, list) else[] print("Ingredients:", ingredients) else: print("No ingredients found in the HTML document") directions_element = soup.find('script', {'type': 'application/ld+json', 'title': 'recipeInstructions'}) if directions_element: directions_data = json.loads(directions_element.string) directions = directions_data if isinstance(directions_data, list) else[] print("Directions:", directions) else: print("No directions found in the HTML document") else: print("Failed to retrieve the page. Status code:", response.status_code) url = "https://www.allrecipes.com/recipe/276505/grandmas-hash-brown-casserole/" scrape_recipe_data(url)
输出结果:
Title: Grandma's Hash Brown Casserole Recipe No ingredients found in the HTML document No directions found in the HTML document Process finished with exit code 0
解决方案
你的问题出在对页面结构化数据的定位错误:allrecipes的食谱数据并非分散在带title属性的独立script标签中,而是全部包含在一个类型为application/ld+json、内容符合Recipe Schema的脚本里。
修改后的代码如下:
import requests from bs4 import BeautifulSoup import json def scrape_recipe_data(url): response = requests.get(url) if response.status_code == 200: soup = BeautifulSoup(response.text, 'html.parser') # 提取页面标题 title_tag = soup.title title_text = title_tag.text if title_tag else "未找到标题" print("标题:", title_text) # 定位包含Recipe schema的LD+JSON脚本 recipe_script = soup.find('script', {'type': 'application/ld+json'}) if recipe_script: try: recipe_data = json.loads(recipe_script.string) # 处理可能的数组格式(部分页面的LD+JSON是数组包裹) if isinstance(recipe_data, list): recipe_data = recipe_data[0] # 提取食材列表 ingredients = recipe_data.get('recipeIngredient', []) print("\n食材:") for idx, ing in enumerate(ingredients, 1): print(f"{idx}. {ing}") # 提取烹饪步骤 instructions = recipe_data.get('recipeInstructions', []) print("\n烹饪步骤:") for idx, step in enumerate(instructions, 1): # 步骤内容可能在'text'或'description'字段中 step_text = step.get('text') or step.get('description', '未找到步骤内容') print(f"{idx}. {step_text}") except json.JSONDecodeError: print("解析结构化数据失败") else: print("未找到食谱结构化数据") else: print(f"页面请求失败,状态码: {response.status_code}") url = "https://www.allrecipes.com/recipe/276505/grandmas-hash-brown-casserole/" scrape_recipe_data(url)
关键修改说明:
- 不再寻找带
title属性的独立脚本,而是直接定位页面中唯一的application/ld+json脚本(部分页面可能是数组,需取第一个元素) - 从解析后的JSON数据中,通过
recipeIngredient字段提取食材列表 - 通过
recipeInstructions字段提取步骤数据,步骤内容通常存在于text或description字段中 - 添加异常处理,避免JSON解析失败导致程序崩溃
运行修改后的代码,会输出完整的食材和烹饪步骤列表。
内容的提问来源于stack exchange,提问作者DeskPrisoner
相关产品推荐
相关产品推荐

