You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python BeautifulSoup提取<br>标签后分组文本并导出至CSV?

针对HTML分组提取与CSV导出的解决方案

核心逻辑是以连续两个<br>作为分组分隔符,逐个提取每个分组内的标题、可选描述和属性,最终导出为CSV。以下是完整步骤和代码:

步骤拆解

  1. 解析HTML文件,定位到<body>节点
  2. 遍历<body>子节点,用连续<br>分割出独立分组
  3. 对每个分组,提取<b>标签的标题、<pre>标签的描述(如果存在),以及剩余的属性文本
  4. 将整理好的数据写入CSV文件

完整代码示例

from bs4 import BeautifulSoup
import csv

# 加载并解析HTML文件(替换为你的文件路径)
with open('your_large_html.html', 'r', encoding='utf-8') as html_file:
    soup = BeautifulSoup(html_file, 'html.parser')

body_content = soup.body
groups = []
current_group = []

# 遍历节点,按连续两个<br>分割分组
for node in body_content.children:
    # 处理纯文本节点,过滤无效空白
    if isinstance(node, str):
        cleaned_text = node.strip()
        if cleaned_text:
            current_group.append(('property', cleaned_text))
        continue
    
    # 检测连续两个<br>,触发分组分割
    if node.name == 'br':
        # 跳过中间空白文本,找下一个有效节点
        next_node = node.next_sibling
        while next_node and isinstance(next_node, str) and not next_node.strip():
            next_node = next_node.next_sibling
        
        if next_node and next_node.name == 'br':
            if current_group:
                groups.append(current_group)
                current_group = []
            continue
    
    # 提取分组内的特定标签内容
    if node.name == 'b':
        current_group.append(('title', node.get_text(strip=True)))
    elif node.name == 'pre':
        current_group.append(('description', node.get_text(strip=True)))

# 处理最后一个未被分隔的分组
if current_group:
    groups.append(current_group)

# 整理数据为字典格式,适配CSV导出
processed_data = []
for group in groups:
    group_dict = {'title': '', 'description': '', 'property': ''}
    for item in group:
        key, value = item
        group_dict[key] = value
    processed_data.append(group_dict)

# 写入CSV文件
with open('groups_result.csv', 'w', newline='', encoding='utf-8') as csv_file:
    writer = csv.DictWriter(csv_file, fieldnames=['title', 'description', 'property'])
    writer.writeheader()
    writer.writerows(processed_data)

关键细节说明

  • 空白处理:所有文本都会用strip()清理多余换行和空格,避免CSV出现无效空白内容
  • 可选描述兼容:如果分组没有<pre>标签,description字段会自动留空,不影响导出
  • 大文件适配:使用Python内置的html.parser解析器,内存占用更低,适合处理大型HTML文件

内容的提问来源于stack exchange,提问作者Steven

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 13:48:23