使用Selenium和BeautifulSoup爬取FoodNetwork食谱链接失败求助
问题描述
我想从https://www.foodnetwork.com/recipes爬取食谱链接,提取每个食谱页的标题、食材和制作步骤并写入recipe_output.txt文件。但提取食谱链接时遇到问题:用soup.find_all("a", class_="m-MediaBlock__a-HeadlineText")获取链接,每次都输出“Found 0 recipe links”,尝试其他选择器也无效。页面中食谱链接的HTML结构如下:
<h3 class="m-MediaBlock__a-Headline"> <a href="//www.foodnetwork.com/recipes/food-network-kitchen/salad-stuffed-peppers-9970168"> <span class="m-MediaBlock__a-HeadlineText">Salad-Stuffed Peppers</span> </a> </h3>
目前recipe_output.txt中没有任何内容,完整代码如下:
options = webdriver.ChromeOptions() options.add_argument("--headless") # Run the browser in headless mode (without a GUI) driver = webdriver.Chrome(options=options) listing_url = "https://www.foodnetwork.com/recipes" driver.get(listing_url) html_content = driver.execute_script("return document.documentElement.outerHTML") soup = BeautifulSoup(html_content, "html.parser") recipe_links = soup.find_all("a", class_="m-MediaBlock__a-HeadlineText") print("Found {} recipe links".format(len(recipe_links))) with open("recipe_output.txt", mode="w", encoding="utf-8") as file: for link in recipe_links: recipe_url = "https:" + link["href"] driver.get(recipe_url) recipe_html_content = driver.execute_script("return document.documentElement.outerHTML") recipe_soup = BeautifulSoup(recipe_html_content, "html.parser") title_element = recipe_soup.find("span", class_="o-AssetTitle__a-HeadlineText") if title_element is not None: recipe_title = title_element.text.strip() file.write("Title: " + recipe_title) ingredient_elements = recipe_soup.find_all("p", class_="o-Ingredients__a-Ingredient") ingredients = [] for ingredient_element in ingredient_elements: ingredient_name_element = ingredient_element.find("span", class_="o-Ingredients__a-Ingredient--CheckboxLabel") if ingredient_name_element is not None: ingredient_name = ingredient_name_element.text.strip() # Clean the ingredient name ingredient_name = re.sub(r"\s*\xa0\s*", " ", ingredient_name) # Remove "xa0" characters if ingredient_name != "Deselect All": ingredients.append(ingredient_name) file.write("Ingredients:\n") for ingredient in ingredients: file.write("- {}\n".format(ingredient)) instructions_elements = recipe_soup.find_all("li", class_="o-Method__m-Step") instructions = [re.sub(r"\s*\xa0\s*", " ", instruction.text.strip()) for instruction in instructions_elements] file.write("Instructions:\n") for i, instruction in enumerate(instructions, start=1): file.write("{}. {}\n".format(i, instruction)) file.write("----------\n") # Quit the webdriver driver.quit()
问题排查与修复方案
核心问题分析
- 选择器匹配错误:
m-MediaBlock__a-HeadlineText是<span>标签的类名,不是<a>标签的。原代码用soup.find_all("a", class_="m-MediaBlock__a-HeadlineText"),相当于找带有这个类的<a>标签,实际该类属于<span>,嵌套在<a>内部,所以匹配不到任何元素。 - 页面加载不充分:headless模式下,可能页面还未完全渲染就获取了HTML,导致目标元素未加载出来。
- headless模式被检测:部分网站会识别headless浏览器并限制内容加载。
修复步骤
1. 修正链接选择器
把查找链接的代码改成以下两种方式之一:
- 方式一:先找到带目标类的
<span>,再获取其父级<a>标签
recipe_span_elements = soup.find_all("span", class_="m-MediaBlock__a-HeadlineText") recipe_links = [span.parent for span in recipe_span_elements if span.parent.name == 'a' and 'href' in span.parent.attrs]
- 方式二:直接通过父级
<h3>定位<a>标签
recipe_links = soup.find_all("a", href=True, parent={"class": "m-MediaBlock__a-Headline"})
2. 添加页面等待机制
在driver.get(listing_url)之后,添加显式等待确保元素加载完成:
from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 最长等待10秒,直到目标元素出现 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "m-MediaBlock__a-HeadlineText")) )
3. 优化headless模式配置
添加参数避免被网站识别为自动化浏览器:
options.add_argument("--disable-blink-features=AutomationControlled") options.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")
完整修复后的代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import re options = webdriver.ChromeOptions() options.add_argument("--headless") options.add_argument("--disable-blink-features=AutomationControlled") options.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") driver = webdriver.Chrome(options=options) listing_url = "https://www.foodnetwork.com/recipes" driver.get(listing_url) # 等待列表页元素加载 try: WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "m-MediaBlock__a-HeadlineText")) ) except: print("页面加载超时,未找到目标元素") driver.quit() exit() html_content = driver.execute_script("return document.documentElement.outerHTML") soup = BeautifulSoup(html_content, "html.parser") # 正确获取食谱链接 recipe_span_elements = soup.find_all("span", class_="m-MediaBlock__a-HeadlineText") recipe_links = [span.parent for span in recipe_span_elements if span.parent.name == 'a' and 'href' in span.parent.attrs] print("Found {} recipe links".format(len(recipe_links))) with open("recipe_output.txt", mode="w", encoding="utf-8") as file: for link in recipe_links: # 处理链接格式,避免重复添加https: recipe_url = "https:" + link["href"] if link["href"].startswith("//") else link["href"] driver.get(recipe_url) # 等待食谱页面加载 try: WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "o-AssetTitle__a-HeadlineText")) ) except: print(f"食谱页面 {recipe_url} 加载超时,跳过") continue recipe_html_content = driver.execute_script("return document.documentElement.outerHTML") recipe_soup = BeautifulSoup(recipe_html_content, "html.parser") # 提取标题 title_element = recipe_soup.find("span", class_="o-AssetTitle__a-HeadlineText") if title_element is not None: recipe_title = title_element.text.strip() file.write(f"Title: {recipe_title}\n") else: file.write("Title: 未获取到标题\n") # 提取食材 ingredient_elements = recipe_soup.find_all("p", class_="o-Ingredients__a-Ingredient") ingredients = [] for ingredient_element in ingredient_elements: ingredient_name_element = ingredient_element.find("span", class_="o-Ingredients__a-Ingredient--CheckboxLabel") if ingredient_name_element is not None: ingredient_name = ingredient_name_element.text.strip() ingredient_name = re.sub(r"\s*\xa0\s*", " ", ingredient_name) if ingredient_name != "Deselect All": ingredients.append(ingredient_name) file.write("Ingredients:\n") if ingredients: for ingredient in ingredients: file.write(f"- {ingredient}\n") else: file.write("- 未获取到食材\n") # 提取制作步骤 instructions_elements = recipe_soup.find_all("li", class_="o-Method__m-Step") instructions = [re.sub(r"\s*\xa0\s*", " ", instruction.text.strip()) for instruction in instructions_elements] file.write("Instructions:\n") if instructions: for i, instruction in enumerate(instructions, start=1): file.write(f"{i}. {instruction}\n") else: file.write("- 未获取到制作步骤\n") file.write("----------\n") driver.quit()
内容的提问来源于stack exchange,提问作者jdavis29
相关产品推荐
相关产品推荐

