You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup仅获取到后半段HTML?技术求助

解决方法

1. 关于打印doc仅显示后半段HTML的问题

print(doc)默认会截断长内容,并不代表BeautifulSoup没有解析完整的HTML。你可以通过以下方式确认:

  • 用print(len(doc))查看解析后的对象长度,对比len(result.text),如果数值接近,说明解析完整。
  • 用print(doc.prettify())输出格式化后的完整HTML,不会被截断。

2. 关于doc.find(text="recipeIngredient")返回None的问题

recipeIngredient通常是食谱页面**结构化数据(JSON-LD)**里的键,或者是HTML元素的属性/类名,不是页面中直接显示的文本节点,所以用find(text=...)找不到。根据数据存储位置,有两种常见提取方式:

方式一:从JSON-LD结构化数据中提取

大多数正规食谱网站会用Schema.org的结构化数据,把食材信息放在<script type="application/ld+json">标签里:

import requests
from bs4 import BeautifulSoup
import json

# 替换为你的目标食谱URL
url = "https://example.com/recipe"
result = requests.get(url)
doc = BeautifulSoup(result.text, "html.parser")

# 定位JSON-LD脚本标签
script_tag = doc.find("script", type="application/ld+json")
if script_tag:
    # 解析JSON内容
    recipe_json = json.loads(script_tag.string)
    # 提取食材列表
    ingredients = recipe_json.get("recipeIngredient", [])
    for ing in ingredients:
        print(ing)

方式二:从HTML元素中直接提取

如果食材是直接写在页面的列表里(比如class为ingredient的<li>标签),可以通过CSS选择器定位:

# 示例:假设食材在class为"ingredients-group"下的li标签中
ingredient_elements = doc.select(".ingredients-group li")
for elem in ingredient_elements:
    # 提取并清理文本
    print(elem.get_text(strip=True))

额外提示

如果用默认的html.parser解析有异常,可以尝试更换解析器,比如lxml或html5lib(需要先通过pip install lxml/pip install html5lib安装):

doc = BeautifulSoup(result.text, "lxml")

内容的提问来源于stack exchange,提问作者iFallOffStuff

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 09:41:29