You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取Allrecipes食谱:食材与步骤解析失败求助

问题

我正在编写程序从allrecipes.com的食谱页面提取标题、URL、食材和烹饪步骤,目前先针对单个页面测试,但食材和步骤的解析模块有问题——HTML中这两部分是列表格式,我无法将其解析为可用数据。作为BS4新手,附上代码及输出结果寻求帮助:

原代码:

import requests
from bs4 import BeautifulSoup
import json


def scrape_recipe_data(url):
    # Make an HTTP request to the provided URL

    response = requests.get(url)

    # Check if the request was successful (status code 200)
    if response.status_code == 200:
        soup = BeautifulSoup(response.text, 'html.parser')

        # Extract title
        title_tag = soup.title
        title_text = title_tag.text if title_tag else "No title found in the HTML document"
        print("Title:", title_text)

        ingredients_element = soup.find('script', {'type': 'application/ld+json', 'title': 'recipeIngredient'}) 

        if ingredients_element:
            ingredients_data = json.loads(ingredients_element.string)
            ingredients = ingredients_data if isinstance(ingredients_data, list) else[]
            print("Ingredients:", ingredients)
        else:
            print("No ingredients found in the HTML document")
        directions_element = soup.find('script', {'type': 'application/ld+json', 'title': 'recipeInstructions'})  

        if directions_element:
            directions_data = json.loads(directions_element.string)
            directions = directions_data if isinstance(directions_data, list) else[]
            print("Directions:", directions)
        else:
            print("No directions found in the HTML document")


    else:
        print("Failed to retrieve the page. Status code:", response.status_code)


url = "https://www.allrecipes.com/recipe/276505/grandmas-hash-brown-casserole/"
scrape_recipe_data(url)

输出结果:

Title: Grandma's Hash Brown Casserole Recipe
No ingredients found in the HTML document
No directions found in the HTML document

Process finished with exit code 0
解决方案

你的问题出在对页面结构化数据的定位错误:allrecipes的食谱数据并非分散在带title属性的独立script标签中,而是全部包含在一个类型为application/ld+json、内容符合Recipe Schema的脚本里。

修改后的代码如下:

import requests
from bs4 import BeautifulSoup
import json


def scrape_recipe_data(url):
    response = requests.get(url)
    if response.status_code == 200:
        soup = BeautifulSoup(response.text, 'html.parser')
        
        # 提取页面标题
        title_tag = soup.title
        title_text = title_tag.text if title_tag else "未找到标题"
        print("标题:", title_text)
        
        # 定位包含Recipe schema的LD+JSON脚本
        recipe_script = soup.find('script', {'type': 'application/ld+json'})
        if recipe_script:
            try:
                recipe_data = json.loads(recipe_script.string)
                # 处理可能的数组格式(部分页面的LD+JSON是数组包裹)
                if isinstance(recipe_data, list):
                    recipe_data = recipe_data[0]
                
                # 提取食材列表
                ingredients = recipe_data.get('recipeIngredient', [])
                print("\n食材:")
                for idx, ing in enumerate(ingredients, 1):
                    print(f"{idx}. {ing}")
                
                # 提取烹饪步骤
                instructions = recipe_data.get('recipeInstructions', [])
                print("\n烹饪步骤:")
                for idx, step in enumerate(instructions, 1):
                    # 步骤内容可能在'text'或'description'字段中
                    step_text = step.get('text') or step.get('description', '未找到步骤内容')
                    print(f"{idx}. {step_text}")
            
            except json.JSONDecodeError:
                print("解析结构化数据失败")
        else:
            print("未找到食谱结构化数据")
    else:
        print(f"页面请求失败,状态码: {response.status_code}")


url = "https://www.allrecipes.com/recipe/276505/grandmas-hash-brown-casserole/"
scrape_recipe_data(url)

关键修改说明:

  • 不再寻找带title属性的独立脚本,而是直接定位页面中唯一的application/ld+json脚本(部分页面可能是数组,需取第一个元素)
  • 从解析后的JSON数据中,通过recipeIngredient字段提取食材列表
  • 通过recipeInstructions字段提取步骤数据,步骤内容通常存在于text或description字段中
  • 添加异常处理,避免JSON解析失败导致程序崩溃

运行修改后的代码,会输出完整的食材和烹饪步骤列表。

内容的提问来源于stack exchange,提问作者DeskPrisoner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 10:45:02