如何用BeautifulSoup更优雅地提取特定HTML标签生成新页面?
更优雅的BeautifulSoup提取特定标签生成HTML文件方案
你可以直接利用BeautifulSoup的DOM操作API构建新HTML文档,无需手动拼接字符串,代码更易维护也更贴合工具设计逻辑:
from bs4 import BeautifulSoup # 读取原HTML文件 with open('P:/Test.html', 'r') as f: soup = BeautifulSoup(f.read(), 'html.parser') # 创建新的HTML文档骨架 new_soup = BeautifulSoup("<html><body></body></html>", 'html.parser') body_tag = new_soup.body # 定义需要保留的目标元素列表 target_elements = [ soup.find('title'), soup.find('p', attrs={'class': 'm-b-0'}), soup.find('div', attrs={'id': 'right-col'}) ] # 将有效元素添加到新文档的body中 for elem in target_elements: if elem: # 过滤找不到的元素,避免写入无效内容 body_tag.append(elem) # 写入输出文件 with open("output1.html", "w") as file: file.write(new_soup.prettify()) # prettify()自动格式化HTML,也可改用str(new_soup)输出紧凑格式
方案优势:
- 完全依托BeautifulSoup原生DOM操作,避免手动拼接字符串可能引发的格式错误
- 通过
append()方法管理元素,更符合HTML文档的结构逻辑 prettify()可自动整理输出格式,提升可读性- 增加元素存在性判断,避免因找不到目标标签导致的异常
内容的提问来源于stack exchange,提问作者Villard
相关产品推荐
相关产品推荐

