You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何扩展Web Scraping代码提取含<strong>的<p>标签食谱步骤?

Extending Your Recipe Scraper to Extract Method Steps

Great job getting the ingredient extraction up and running! Let's break down how to expand your code to pull the recipe method steps exactly as you need them—focusing only on <p> tags that contain <strong> elements within the correct <div> container.

Step 1: Fix & Refine the Target Container

First, let's update the div selection to match your research: we need the <div> with both id="sr-recipe-callout" and class="recipe-callout". This ensures we're targeting the exact section of the page with the method steps, avoiding any accidental matches elsewhere.

Step 2: Add Step Extraction Logic

To filter only <p> tags that include a <strong> element, we can use a conditional check in a list comprehension. The find() method returns None if no matching element is found, so we can use that as a simple boolean filter to keep only the relevant steps.

Full Updated Code

Here's the evolved version of your code, with comments explaining each new part:

import requests
from bs4 import BeautifulSoup
from scraper_api import ScraperAPIClient
from splinter import Browser
from webdriver_manager.chrome import ChromeDriverManager

# Keeping your Splinter/Browser setup in case you need it for dynamic content later
executable_path = {'executable_path': ChromeDriverManager().install()}
browser = Browser('chrome', **executable_path)

# Fetch the recipe page
resp = requests.get("https://www.simplyrecipes.com/recipes/cooking_for_two_strawberry_almond_oat_smoothie/")
soup = BeautifulSoup(resp.text, "html.parser")

# Target the exact div (matches your research: id + class combination)
div_ = soup.find("div", attrs={"id": "sr-recipe-callout", "class": "recipe-callout"})

# Extract recipe name (stored in a variable for clearer code)
recipe_name = "_".join(div_.find("h2").text.split())

# Your original working ingredient extraction
ingredients = [x.text for x in div_.findAll("li", attrs={"class": "ingredient"})]

# Extract method steps: only keep <p> tags that contain a <strong> element
steps = [p.text.strip() for p in div_.findAll("p") if p.find("strong")]

# Build a structured recipe dictionary with both ingredients and steps
recipes = {
    recipe_name: {
        "ingredients": ingredients,
        "method_steps": steps
    }
}

print(recipes)

Key Explanations

  • Precise Div Selection: By including both id and class in the find() call, we eliminate ambiguity and ensure we're working with the right section of the page.
  • Step Filtering: The condition if p.find("strong") acts as a gatekeeper—only <p> tags that have a <strong> child element make it into the steps list.
  • Structured Data: We've updated the recipes dictionary to separate ingredients and steps into distinct keys, making the data much easier to parse, store, or use later in your project.

Example Output

You'll get a clean, structured dictionary like this (abbreviated for brevity):

{
    'Strawberry_Almond_Oat_Smoothie_Recipe': {
        'ingredients': ['1/2 cup uncooked old-fashioned rolled oats', '2 cups frozen strawberries', ...],
        'method_steps': [
            '1. Place the oats in a blender and blend for 30 seconds to turn them into a coarse flour.',
            '2. Add the strawberries, yogurt, almond milk, banana, almond extract, and honey (if using) to the blender.',
            ...
        ]
    }
}

内容的提问来源于stack exchange,提问作者William

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 19:37:43