You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Jupyter中使用Python合并多个HTML文件为一个?

合并多个HTML文件为一个的实现思路(Jupyter环境)

和处理Excel结构化数据的逻辑不同,HTML属于文档型内容,需要针对其文本或DOM结构来实现合并,以下是几种实用思路:

1. 直接读取文本拼接(快速简单)

适合不需要复杂结构处理的场景,比如合并多个HTML片段,或提取完整HTML的主体内容进行拼接:

import os

data_folder = 'C:\\Users\\hhhh\\Desktop\\test'
merged_parts = []

# 初始化合并后的HTML基础结构
merged_parts.append('<!DOCTYPE html><html><head><title>合并结果</title></head><body>')

for file in os.listdir(data_folder):
    if file.endswith('.html'):
        print(f'加载文件 {file}...')
        file_path = os.path.join(data_folder, file)
        with open(file_path, 'r', encoding='utf-8') as f:
            content = f.read()
            # 提取<body>标签内的内容,避免重复的根结构
            start_pos = content.find('<body>') + len('<body>')
            end_pos = content.rfind('</body>')
            if start_pos != -1 and end_pos != -1:
                content = content[start_pos:end_pos]
            merged_parts.append(content)

# 补充HTML尾部
merged_parts.append('</body></html>')

# 保存合并后的文件
with open(os.path.join(data_folder, 'merged.html'), 'w', encoding='utf-8') as f:
    f.write('\n'.join(merged_parts))

2. 用BeautifulSoup结构化合并(安全可靠)

如果需要严格保证HTML结构的合法性,避免标签嵌套错误,可以用BeautifulSoup处理DOM树:

from bs4 import BeautifulSoup
import os

data_folder = 'C:\\Users\\hhhh\\Desktop\\test'
# 创建一个空的基础HTML结构
merged_soup = BeautifulSoup('<html><head><title>合并结果</title></head><body></body></html>', 'html.parser')
target_body = merged_soup.body

for file in os.listdir(data_folder):
    if file.endswith('.html'):
        print(f'加载文件 {file}...')
        file_path = os.path.join(data_folder, file)
        with open(file_path, 'r', encoding='utf-8') as f:
            current_soup = BeautifulSoup(f.read(), 'html.parser')
            # 将当前文件<body>内的所有元素添加到合并后的<body>中
            for element in current_soup.body.contents:
                if element.strip():  # 跳过空内容节点
                    target_body.append(element)

# 保存结果
with open(os.path.join(data_folder, 'merged_soup.html'), 'w', encoding='utf-8') as f:
    f.write(str(merged_soup))

3. 基于模板的合并(规整美观)

如果需要给合并后的页面添加统一的样式、导航或章节划分,可以用Jinja2模板引擎:
先在Jupyter中安装依赖:

!pip install jinja2

然后执行合并代码:

from jinja2 import Template
import os

data_folder = 'C:\\Users\\hhhh\\Desktop\\test'
section_list = []

for file in os.listdir(data_folder):
    if file.endswith('.html'):
        print(f'加载文件 {file}...')
        file_path = os.path.join(data_folder, file)
        with open(file_path, 'r', encoding='utf-8') as f:
            content = f.read()
            # 提取<body>内的内容
            start_pos = content.find('<body>') + len('<body>')
            end_pos = content.rfind('</body>')
            if start_pos != -1 and end_pos != -1:
                content = content[start_pos:end_pos]
            section_list.append({
                'title': file.replace('.html', ''),
                'content': content
            })

# 定义合并模板,可自定义样式和结构
template = Template('''
<!DOCTYPE html>
<html>
<head>
    <title>合并文档</title>
    <style>
        .section { margin: 2rem auto; padding: 1rem; max-width: 1000px; border-bottom: 1px solid #eee; }
        .section h2 { color: #2c3e50; }
    </style>
</head>
<body>
    {% for sec in section_list %}
    <div class="section">
        <h2>{{ sec.title }}</h2>
        {{ sec.content }}
    </div>
    {% endfor %}
</body>
</html>
''')

# 渲染模板并保存
merged_html = template.render(section_list=section_list)
with open(os.path.join(data_folder, 'merged_template.html'), 'w', encoding='utf-8') as f:
    f.write(merged_html)

和Excel处理逻辑的区别

你之前处理Excel的代码是读取结构化表格数据,将多个DataFrame拼接;而HTML是非结构化的文档内容,无法直接用类似pd.read_excel的方式读取后拼接,必须针对文本或DOM结构做处理。

内容的提问来源于stack exchange,提问作者Maha Mohammed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 18:42:40