You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup更优雅地提取特定HTML标签生成新页面?

更优雅的BeautifulSoup提取特定标签生成HTML文件方案

你可以直接利用BeautifulSoup的DOM操作API构建新HTML文档,无需手动拼接字符串,代码更易维护也更贴合工具设计逻辑:

from bs4 import BeautifulSoup

# 读取原HTML文件
with open('P:/Test.html', 'r') as f:
    soup = BeautifulSoup(f.read(), 'html.parser')

# 创建新的HTML文档骨架
new_soup = BeautifulSoup("<html><body></body></html>", 'html.parser')
body_tag = new_soup.body

# 定义需要保留的目标元素列表
target_elements = [
    soup.find('title'),
    soup.find('p', attrs={'class': 'm-b-0'}),
    soup.find('div', attrs={'id': 'right-col'})
]

# 将有效元素添加到新文档的body中
for elem in target_elements:
    if elem:  # 过滤找不到的元素,避免写入无效内容
        body_tag.append(elem)

# 写入输出文件
with open("output1.html", "w") as file:
    file.write(new_soup.prettify())  # prettify()自动格式化HTML,也可改用str(new_soup)输出紧凑格式

方案优势:

  • 完全依托BeautifulSoup原生DOM操作,避免手动拼接字符串可能引发的格式错误
  • 通过append()方法管理元素,更符合HTML文档的结构逻辑
  • prettify()可自动整理输出格式,提升可读性
  • 增加元素存在性判断,避免因找不到目标标签导致的异常

内容的提问来源于stack exchange,提问作者Villard

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 04:35:09