You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium和BeautifulSoup爬取FoodNetwork食谱链接失败求助

问题描述

我想从https://www.foodnetwork.com/recipes爬取食谱链接,提取每个食谱页的标题、食材和制作步骤并写入recipe_output.txt文件。但提取食谱链接时遇到问题:用soup.find_all("a", class_="m-MediaBlock__a-HeadlineText")获取链接,每次都输出“Found 0 recipe links”,尝试其他选择器也无效。页面中食谱链接的HTML结构如下:

<h3 class="m-MediaBlock__a-Headline">
    <a href="//www.foodnetwork.com/recipes/food-network-kitchen/salad-stuffed-peppers-9970168">
      <span class="m-MediaBlock__a-HeadlineText">Salad-Stuffed Peppers</span>
      
    </a>
  </h3>

目前recipe_output.txt中没有任何内容,完整代码如下:

options = webdriver.ChromeOptions()
options.add_argument("--headless")  # Run the browser in headless mode (without a GUI)
driver = webdriver.Chrome(options=options)

listing_url = "https://www.foodnetwork.com/recipes"
driver.get(listing_url)

html_content = driver.execute_script("return document.documentElement.outerHTML")
soup = BeautifulSoup(html_content, "html.parser")

recipe_links = soup.find_all("a", class_="m-MediaBlock__a-HeadlineText")
print("Found {} recipe links".format(len(recipe_links)))

with open("recipe_output.txt", mode="w", encoding="utf-8") as file:
    for link in recipe_links:
        recipe_url = "https:" + link["href"]
        driver.get(recipe_url)
        recipe_html_content = driver.execute_script("return document.documentElement.outerHTML")
        recipe_soup = BeautifulSoup(recipe_html_content, "html.parser")
        
        title_element = recipe_soup.find("span", class_="o-AssetTitle__a-HeadlineText")
        if title_element is not None:
            recipe_title = title_element.text.strip()
            file.write("Title: " + recipe_title)
            
        ingredient_elements = recipe_soup.find_all("p", class_="o-Ingredients__a-Ingredient")
        ingredients = []
        for ingredient_element in ingredient_elements:
            ingredient_name_element = ingredient_element.find("span", class_="o-Ingredients__a-Ingredient--CheckboxLabel")
            if ingredient_name_element is not None:
                ingredient_name = ingredient_name_element.text.strip()
                
                # Clean the ingredient name
                ingredient_name = re.sub(r"\s*\xa0\s*", " ", ingredient_name)  # Remove "xa0" characters
                if ingredient_name != "Deselect All":
                    ingredients.append(ingredient_name)
                    
        file.write("Ingredients:\n")
        for ingredient in ingredients:
            file.write("- {}\n".format(ingredient))

        instructions_elements = recipe_soup.find_all("li", class_="o-Method__m-Step")
        instructions = [re.sub(r"\s*\xa0\s*", " ", instruction.text.strip()) for instruction in instructions_elements]
        file.write("Instructions:\n")
        for i, instruction in enumerate(instructions, start=1):
            file.write("{}. {}\n".format(i, instruction))
            
        file.write("----------\n")

# Quit the webdriver
driver.quit()
问题排查与修复方案

核心问题分析

  1. 选择器匹配错误:m-MediaBlock__a-HeadlineText是<span>标签的类名,不是<a>标签的。原代码用soup.find_all("a", class_="m-MediaBlock__a-HeadlineText"),相当于找带有这个类的<a>标签,实际该类属于<span>,嵌套在<a>内部,所以匹配不到任何元素。
  2. 页面加载不充分:headless模式下,可能页面还未完全渲染就获取了HTML,导致目标元素未加载出来。
  3. headless模式被检测:部分网站会识别headless浏览器并限制内容加载。

修复步骤

1. 修正链接选择器

把查找链接的代码改成以下两种方式之一:

  • 方式一:先找到带目标类的<span>,再获取其父级<a>标签
recipe_span_elements = soup.find_all("span", class_="m-MediaBlock__a-HeadlineText")
recipe_links = [span.parent for span in recipe_span_elements if span.parent.name == 'a' and 'href' in span.parent.attrs]
  • 方式二:直接通过父级<h3>定位<a>标签
recipe_links = soup.find_all("a", href=True, parent={"class": "m-MediaBlock__a-Headline"})

2. 添加页面等待机制

在driver.get(listing_url)之后,添加显式等待确保元素加载完成:

from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 最长等待10秒,直到目标元素出现
WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.CLASS_NAME, "m-MediaBlock__a-HeadlineText"))
)

3. 优化headless模式配置

添加参数避免被网站识别为自动化浏览器:

options.add_argument("--disable-blink-features=AutomationControlled")
options.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")

完整修复后的代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import re

options = webdriver.ChromeOptions()
options.add_argument("--headless")
options.add_argument("--disable-blink-features=AutomationControlled")
options.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")
driver = webdriver.Chrome(options=options)

listing_url = "https://www.foodnetwork.com/recipes"
driver.get(listing_url)

# 等待列表页元素加载
try:
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "m-MediaBlock__a-HeadlineText"))
    )
except:
    print("页面加载超时,未找到目标元素")
    driver.quit()
    exit()

html_content = driver.execute_script("return document.documentElement.outerHTML")
soup = BeautifulSoup(html_content, "html.parser")

# 正确获取食谱链接
recipe_span_elements = soup.find_all("span", class_="m-MediaBlock__a-HeadlineText")
recipe_links = [span.parent for span in recipe_span_elements if span.parent.name == 'a' and 'href' in span.parent.attrs]

print("Found {} recipe links".format(len(recipe_links)))

with open("recipe_output.txt", mode="w", encoding="utf-8") as file:
    for link in recipe_links:
        # 处理链接格式,避免重复添加https:
        recipe_url = "https:" + link["href"] if link["href"].startswith("//") else link["href"]
        driver.get(recipe_url)
        
        # 等待食谱页面加载
        try:
            WebDriverWait(driver, 10).until(
                EC.presence_of_element_located((By.CLASS_NAME, "o-AssetTitle__a-HeadlineText"))
            )
        except:
            print(f"食谱页面 {recipe_url} 加载超时,跳过")
            continue
            
        recipe_html_content = driver.execute_script("return document.documentElement.outerHTML")
        recipe_soup = BeautifulSoup(recipe_html_content, "html.parser")
        
        # 提取标题
        title_element = recipe_soup.find("span", class_="o-AssetTitle__a-HeadlineText")
        if title_element is not None:
            recipe_title = title_element.text.strip()
            file.write(f"Title: {recipe_title}\n")
        else:
            file.write("Title: 未获取到标题\n")
            
        # 提取食材
        ingredient_elements = recipe_soup.find_all("p", class_="o-Ingredients__a-Ingredient")
        ingredients = []
        for ingredient_element in ingredient_elements:
            ingredient_name_element = ingredient_element.find("span", class_="o-Ingredients__a-Ingredient--CheckboxLabel")
            if ingredient_name_element is not None:
                ingredient_name = ingredient_name_element.text.strip()
                ingredient_name = re.sub(r"\s*\xa0\s*", " ", ingredient_name)
                if ingredient_name != "Deselect All":
                    ingredients.append(ingredient_name)
                    
        file.write("Ingredients:\n")
        if ingredients:
            for ingredient in ingredients:
                file.write(f"- {ingredient}\n")
        else:
            file.write("- 未获取到食材\n")

        # 提取制作步骤
        instructions_elements = recipe_soup.find_all("li", class_="o-Method__m-Step")
        instructions = [re.sub(r"\s*\xa0\s*", " ", instruction.text.strip()) for instruction in instructions_elements]
        file.write("Instructions:\n")
        if instructions:
            for i, instruction in enumerate(instructions, start=1):
                file.write(f"{i}. {instruction}\n")
        else:
            file.write("- 未获取到制作步骤\n")
            
        file.write("----------\n")

driver.quit()

内容的提问来源于stack exchange,提问作者jdavis29

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 00:43:09