如何扩展Web Scraping代码提取含<strong>的<p>标签食谱步骤?
Great job getting the ingredient extraction up and running! Let's break down how to expand your code to pull the recipe method steps exactly as you need them—focusing only on <p> tags that contain <strong> elements within the correct <div> container.
Step 1: Fix & Refine the Target Container
First, let's update the div selection to match your research: we need the <div> with both id="sr-recipe-callout" and class="recipe-callout". This ensures we're targeting the exact section of the page with the method steps, avoiding any accidental matches elsewhere.
Step 2: Add Step Extraction Logic
To filter only <p> tags that include a <strong> element, we can use a conditional check in a list comprehension. The find() method returns None if no matching element is found, so we can use that as a simple boolean filter to keep only the relevant steps.
Full Updated Code
Here's the evolved version of your code, with comments explaining each new part:
import requests from bs4 import BeautifulSoup from scraper_api import ScraperAPIClient from splinter import Browser from webdriver_manager.chrome import ChromeDriverManager # Keeping your Splinter/Browser setup in case you need it for dynamic content later executable_path = {'executable_path': ChromeDriverManager().install()} browser = Browser('chrome', **executable_path) # Fetch the recipe page resp = requests.get("https://www.simplyrecipes.com/recipes/cooking_for_two_strawberry_almond_oat_smoothie/") soup = BeautifulSoup(resp.text, "html.parser") # Target the exact div (matches your research: id + class combination) div_ = soup.find("div", attrs={"id": "sr-recipe-callout", "class": "recipe-callout"}) # Extract recipe name (stored in a variable for clearer code) recipe_name = "_".join(div_.find("h2").text.split()) # Your original working ingredient extraction ingredients = [x.text for x in div_.findAll("li", attrs={"class": "ingredient"})] # Extract method steps: only keep <p> tags that contain a <strong> element steps = [p.text.strip() for p in div_.findAll("p") if p.find("strong")] # Build a structured recipe dictionary with both ingredients and steps recipes = { recipe_name: { "ingredients": ingredients, "method_steps": steps } } print(recipes)
Key Explanations
- Precise Div Selection: By including both
idandclassin thefind()call, we eliminate ambiguity and ensure we're working with the right section of the page. - Step Filtering: The condition
if p.find("strong")acts as a gatekeeper—only<p>tags that have a<strong>child element make it into thestepslist. - Structured Data: We've updated the
recipesdictionary to separate ingredients and steps into distinct keys, making the data much easier to parse, store, or use later in your project.
Example Output
You'll get a clean, structured dictionary like this (abbreviated for brevity):
{ 'Strawberry_Almond_Oat_Smoothie_Recipe': { 'ingredients': ['1/2 cup uncooked old-fashioned rolled oats', '2 cups frozen strawberries', ...], 'method_steps': [ '1. Place the oats in a blender and blend for 30 seconds to turn them into a coarse flour.', '2. Add the strawberries, yogurt, almond milk, banana, almond extract, and honey (if using) to the blender.', ... ] } }
内容的提问来源于stack exchange,提问作者William

