如何在Jupyter中使用Python合并多个HTML文件为一个?
合并多个HTML文件为一个的实现思路(Jupyter环境)
和处理Excel结构化数据的逻辑不同,HTML属于文档型内容,需要针对其文本或DOM结构来实现合并,以下是几种实用思路:
1. 直接读取文本拼接(快速简单)
适合不需要复杂结构处理的场景,比如合并多个HTML片段,或提取完整HTML的主体内容进行拼接:
import os data_folder = 'C:\\Users\\hhhh\\Desktop\\test' merged_parts = [] # 初始化合并后的HTML基础结构 merged_parts.append('<!DOCTYPE html><html><head><title>合并结果</title></head><body>') for file in os.listdir(data_folder): if file.endswith('.html'): print(f'加载文件 {file}...') file_path = os.path.join(data_folder, file) with open(file_path, 'r', encoding='utf-8') as f: content = f.read() # 提取<body>标签内的内容,避免重复的根结构 start_pos = content.find('<body>') + len('<body>') end_pos = content.rfind('</body>') if start_pos != -1 and end_pos != -1: content = content[start_pos:end_pos] merged_parts.append(content) # 补充HTML尾部 merged_parts.append('</body></html>') # 保存合并后的文件 with open(os.path.join(data_folder, 'merged.html'), 'w', encoding='utf-8') as f: f.write('\n'.join(merged_parts))
2. 用BeautifulSoup结构化合并(安全可靠)
如果需要严格保证HTML结构的合法性,避免标签嵌套错误,可以用BeautifulSoup处理DOM树:
from bs4 import BeautifulSoup import os data_folder = 'C:\\Users\\hhhh\\Desktop\\test' # 创建一个空的基础HTML结构 merged_soup = BeautifulSoup('<html><head><title>合并结果</title></head><body></body></html>', 'html.parser') target_body = merged_soup.body for file in os.listdir(data_folder): if file.endswith('.html'): print(f'加载文件 {file}...') file_path = os.path.join(data_folder, file) with open(file_path, 'r', encoding='utf-8') as f: current_soup = BeautifulSoup(f.read(), 'html.parser') # 将当前文件<body>内的所有元素添加到合并后的<body>中 for element in current_soup.body.contents: if element.strip(): # 跳过空内容节点 target_body.append(element) # 保存结果 with open(os.path.join(data_folder, 'merged_soup.html'), 'w', encoding='utf-8') as f: f.write(str(merged_soup))
3. 基于模板的合并(规整美观)
如果需要给合并后的页面添加统一的样式、导航或章节划分,可以用Jinja2模板引擎:
先在Jupyter中安装依赖:
!pip install jinja2
然后执行合并代码:
from jinja2 import Template import os data_folder = 'C:\\Users\\hhhh\\Desktop\\test' section_list = [] for file in os.listdir(data_folder): if file.endswith('.html'): print(f'加载文件 {file}...') file_path = os.path.join(data_folder, file) with open(file_path, 'r', encoding='utf-8') as f: content = f.read() # 提取<body>内的内容 start_pos = content.find('<body>') + len('<body>') end_pos = content.rfind('</body>') if start_pos != -1 and end_pos != -1: content = content[start_pos:end_pos] section_list.append({ 'title': file.replace('.html', ''), 'content': content }) # 定义合并模板,可自定义样式和结构 template = Template(''' <!DOCTYPE html> <html> <head> <title>合并文档</title> <style> .section { margin: 2rem auto; padding: 1rem; max-width: 1000px; border-bottom: 1px solid #eee; } .section h2 { color: #2c3e50; } </style> </head> <body> {% for sec in section_list %} <div class="section"> <h2>{{ sec.title }}</h2> {{ sec.content }} </div> {% endfor %} </body> </html> ''') # 渲染模板并保存 merged_html = template.render(section_list=section_list) with open(os.path.join(data_folder, 'merged_template.html'), 'w', encoding='utf-8') as f: f.write(merged_html)
和Excel处理逻辑的区别
你之前处理Excel的代码是读取结构化表格数据,将多个DataFrame拼接;而HTML是非结构化的文档内容,无法直接用类似pd.read_excel的方式读取后拼接,必须针对文本或DOM结构做处理。
内容的提问来源于stack exchange,提问作者Maha Mohammed
相关产品推荐
相关产品推荐

