You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取<br>标签间无标签文本?实现标题与段落分数组存储

解决方案

核心思路

通过BeautifulSoup遍历节点,先提取所有<strong>标签内的标题,再针对每个标题,收集其后续直到下一个<strong>标签前的所有文本节点(跳过<br>标签),分别存入两个数组。

示例代码实现

假设你的HTML结构类似如下(如果实际结构不同,可微调节点遍历逻辑):

<div class="target-container">
    <strong>产品介绍</strong><br>
    这是一款主打便携的智能设备<br>
    续航可达72小时<br>
    <strong>使用说明</strong><br>
    首次使用请充电3小时<br>
    长按电源键2秒开机
</div>

对应的Python处理代码:

from bs4 import BeautifulSoup
# 如果你是用Selenium获取的页面源码,这里替换成driver.page_source
html_content = """上面的HTML内容"""
soup = BeautifulSoup(html_content, "html.parser")

titles = []
paragraphs = []  # 每个元素对应一个标题下的段落列表

# 获取所有标题节点
title_nodes = soup.find_all("strong")

for title_node in title_nodes:
    # 提取标题文本并加入数组
    title_text = title_node.get_text(strip=True)
    if title_text:
        titles.append(title_text)
    
    # 收集当前标题后的所有段落文本
    current_paragraphs = []
    # 遍历标题节点的后续兄弟节点
    next_sibling = title_node.next_sibling
    while next_sibling:
        # 如果遇到下一个strong标签,停止遍历
        if next_sibling.name == "strong":
            break
        # 如果是文本节点且内容非空,处理后加入列表
        if next_sibling.string and next_sibling.string.strip():
            current_paragraphs.append(next_sibling.string.strip())
        # 移动到下一个兄弟节点
        next_sibling = next_sibling.next_sibling
    # 将当前标题对应的段落列表加入总数组
    paragraphs.append(current_paragraphs)

# 打印结果验证
print("标题数组:", titles)
print("段落数组:", paragraphs)

适配Selenium的场景

如果你的流程是用Selenium打开页面后提取内容,只需将html_content替换为Selenium获取的页面源码:

from selenium import webdriver
from bs4 import BeautifulSoup

driver = webdriver.Chrome()
driver.get("目标页面URL")
html_content = driver.page_source
# 后续处理同上
driver.quit()

注意事项

  • 如果你的HTML中<br>是自闭合标签(<br/>),代码逻辑无需调整,BeautifulSoup会自动识别。
  • 若文本中包含多余的换行或空格,可通过strip()或正则表达式进一步清洗,比如用re.sub(r'\s+', ' ', text)合并连续空白。

内容的提问来源于stack exchange,提问作者yerbaMatte

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 20:15:33