You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬取求助:Allrecipes食材替换数据爬取代码调试

修正后的食材替换信息爬取代码

首先补充必要的依赖导入,然后修正定位逻辑和数据提取方式:

import requests
from bs4 import BeautifulSoup

def scrape_ingredient_substitutions():  
    """Scrape the Allrecipes website for common ingredient substitutions."""  
    url = "https://www.allrecipes.com/article/common-ingredient-substitutions/"  
    response = requests.get(url)
    # 检查请求是否成功
    if response.status_code != 200:
        print(f"请求失败,状态码:{response.status_code}")
        return {}
    
    soup = BeautifulSoup(response.content, "html.parser")  

    substitutions = {}  
    # 获取页面中所有的tr标签,排除表头行
    all_tr = soup.find_all("tr")
    # 从第2个tr开始遍历,第一个是表头
    for tr in all_tr[1:]:
        tds = tr.find_all("td")
        # 确保当前行有3个有效td节点
        if len(tds) == 3:
            ingredient = tds[0].get_text(strip=True)
            amount = tds[1].get_text(strip=True)
            substitution = tds[2].get_text(strip=True)
            substitutions[ingredient] = {"amount": amount, "substitution": substitution}  

    return substitutions  

subs = scrape_ingredient_substitutions()  
# 打印部分结果验证
for ing, info in list(subs.items())[:5]:
    print(f"原食材:{ing},用量:{info['amount']},替换品:{info['substitution']}")

关键修改说明

  • 补充依赖导入:原代码缺少requests和BeautifulSoup的导入语句,运行会直接报错,必须先添加。
  • 修正tr标签定位:
    • 原代码用soup.find("tr", class_=None)仅能获取单个tr标签,改为soup.find_all("tr")获取所有行。
    • 跳过第一个tr(表头行),避免把标题文本混入数据。
  • 正确提取td文本:
    • 每个数据行的tr包含3个td,分别对应「原食材」「用量」「替换品」,通过索引tds[0]、tds[1]、tds[2]分别提取。
    • 使用get_text(strip=True)去除文本前后的空格和换行,让数据更整洁。
  • 增加请求校验:添加状态码检查,避免请求失败后继续执行后续逻辑导致报错。
  • 修复缩进问题:原代码中循环和变量赋值的缩进不符合Python规范,统一调整为4空格缩进。

内容的提问来源于stack exchange,提问作者Kemish Jimenez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 06:47:23