如何按顺序拼接p标签并将table标签独立存入列表?代码问题排查
问题解决:连续P标签拼接与Table标签独立存储
原代码问题分析
- 每次循环都新建
content_dict,导致每个P标签内容无法累积拼接 - 使用
descendants遍历会获取所有嵌套节点,可能引入冗余内容 - 循环结束后未处理剩余的P标签内容
修正后的代码
from bs4 import BeautifulSoup import html2text import json converter = html2text.HTML2Text() soup = BeautifulSoup(data3, 'html.parser') content_items = [] current_p_content = "" # 临时存储连续的P标签内容 title = "35.23.060 - DR Zone Standards" # 按文档顺序遍历所有顶层的p和table标签 for tag in soup.find_all(['p', 'table'], recursive=False): if tag.name == 'p': # 拼接当前P标签的转换后内容 current_p_content += converter.handle(str(tag)) elif tag.name == 'table': # 如果有累积的P内容,先存入列表 if current_p_content.strip(): content_items.append({'Title': title, 'Content': current_p_content.strip()}) current_p_content = "" # 清空临时容器 # 将Table作为独立项存入 content_items.append({'Title': title, 'Content': converter.handle(str(tag))}) # 处理循环结束后剩余的P内容 if current_p_content.strip(): content_items.append({'Title': title, 'Content': current_p_content.strip()}) # 打印结果 print(json.dumps(content_items, indent=4))
关键改进点
- 用
current_p_content临时变量累积连续P标签内容,实现拼接效果 - 使用
find_all(['p', 'table'], recursive=False)获取顶层目标标签,避免嵌套节点干扰 - 循环结束后检查剩余P内容,防止遗漏
- 加入
strip()处理内容前后空白,避免无效空项
内容的提问来源于stack exchange,提问作者Simon Palmer
相关产品推荐
相关产品推荐

