You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取HTML中content与后续标签间的文本?

提取content与后续同类标题块之间的文本(Python实现)

用Python的BeautifulSoup库就能搞定这个需求,步骤和代码如下:

首先安装依赖:

pip install beautifulsoup4

核心思路

  1. 解析HTML内容,定位到包含content:的<strong>标签所在的<p>块
  2. 从这个<p>块开始,遍历后续所有同级<p>元素
  3. 收集每个<p>的文本,直到遇到下一个包含带冒号的标题(比如time:、location:这类)的<p><strong>块为止

示例代码

from bs4 import BeautifulSoup

# 替换成你的实际HTML内容
html = """
<p><strong> name:</strong></p>
<p> sdfsdf </p>
<p><strong> content:</strong></p>
<p> yangben</p>
<p> dsfs </p>
<p> dfsds </p>
<p> sdfs </p>
<p><strong> time:</strong></p>
<p> 2020-10-10</p>
<p><strong> ll:</strong></p>
<p> 2020-10-10</p>
"""

soup = BeautifulSoup(html, 'html.parser')

# 定位content标题块
content_header = soup.find('strong', string=lambda t: t and 'content:' in t.strip())
if not content_header:
    print("找不到content对应的标题块")
else:
    current_p = content_header.parent
    result = []
    
    # 遍历后续兄弟p元素
    for sibling in current_p.find_next_siblings('p'):
        # 检查是否是下一个标题块(带冒号的strong)
        strong = sibling.find('strong')
        if strong and ':' in strong.text.strip():
            break
        # 提取非空文本
        text = sibling.text.strip()
        if text:
            result.append(text)
    
    # 输出结果
    print("提取到的内容:")
    for line in result:
        print(line)

代码说明

  • 用lambda表达式匹配content:,能忽略文本前后的空格,避免匹配失败
  • 判断停止条件时,只要遇到带冒号的<strong>标题就停止,兼容time、location等任何同类标签
  • 自动过滤空文本,只保留有效内容

运行结果

提取到的内容:
yangben
dsfs
dfsds
sdfs

内容的提问来源于stack exchange,提问作者Lei Hao

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 20:32:37