如何用Python BeautifulSoup提取<br>标签后分组文本并导出至CSV?
针对HTML分组提取与CSV导出的解决方案
核心逻辑是以连续两个<br>作为分组分隔符,逐个提取每个分组内的标题、可选描述和属性,最终导出为CSV。以下是完整步骤和代码:
步骤拆解
- 解析HTML文件,定位到
<body>节点 - 遍历
<body>子节点,用连续<br>分割出独立分组 - 对每个分组,提取
<b>标签的标题、<pre>标签的描述(如果存在),以及剩余的属性文本 - 将整理好的数据写入CSV文件
完整代码示例
from bs4 import BeautifulSoup import csv # 加载并解析HTML文件(替换为你的文件路径) with open('your_large_html.html', 'r', encoding='utf-8') as html_file: soup = BeautifulSoup(html_file, 'html.parser') body_content = soup.body groups = [] current_group = [] # 遍历节点,按连续两个<br>分割分组 for node in body_content.children: # 处理纯文本节点,过滤无效空白 if isinstance(node, str): cleaned_text = node.strip() if cleaned_text: current_group.append(('property', cleaned_text)) continue # 检测连续两个<br>,触发分组分割 if node.name == 'br': # 跳过中间空白文本,找下一个有效节点 next_node = node.next_sibling while next_node and isinstance(next_node, str) and not next_node.strip(): next_node = next_node.next_sibling if next_node and next_node.name == 'br': if current_group: groups.append(current_group) current_group = [] continue # 提取分组内的特定标签内容 if node.name == 'b': current_group.append(('title', node.get_text(strip=True))) elif node.name == 'pre': current_group.append(('description', node.get_text(strip=True))) # 处理最后一个未被分隔的分组 if current_group: groups.append(current_group) # 整理数据为字典格式,适配CSV导出 processed_data = [] for group in groups: group_dict = {'title': '', 'description': '', 'property': ''} for item in group: key, value = item group_dict[key] = value processed_data.append(group_dict) # 写入CSV文件 with open('groups_result.csv', 'w', newline='', encoding='utf-8') as csv_file: writer = csv.DictWriter(csv_file, fieldnames=['title', 'description', 'property']) writer.writeheader() writer.writerows(processed_data)
关键细节说明
- 空白处理:所有文本都会用
strip()清理多余换行和空格,避免CSV出现无效空白内容 - 可选描述兼容:如果分组没有
<pre>标签,description字段会自动留空,不影响导出 - 大文件适配:使用Python内置的
html.parser解析器,内存占用更低,适合处理大型HTML文件
内容的提问来源于stack exchange,提问作者Steven
相关产品推荐
相关产品推荐

