You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中合并ODT文件?两种尝试方案均失败

解决方案:合并多个ODT文件为单个ODT/DOCX

方案一:直接操作ODT文件结构合并(纯Python免费方案)

ODT本质是ZIP压缩包,核心内容存储在content.xml中,可通过解压、合并XML内容、重新打包实现合并:

  1. 安装依赖(Python内置zipfile,仅需额外安装XML处理库):
pip install lxml
  1. 合并代码示例:
import zipfile
from lxml import etree
import os
import tempfile

def merge_odts(input_files, output_file):
    # 创建临时目录处理文件
    with tempfile.TemporaryDirectory() as tmpdir:
        # 以第一个文件为基础,解压所有资源
        with zipfile.ZipFile(input_files[0], 'r') as base_zip:
            base_zip.extractall(tmpdir)
        
        # 解析基础文件的content.xml
        content_path = os.path.join(tmpdir, 'content.xml')
        tree = etree.parse(content_path)
        root = tree.getroot()
        ns = {'text': 'urn:oasis:names:tc:opendocument:xmlns:text:1.0'}
        
        # 获取正文节点,用于插入后续内容
        body = root.find('.//text:body', ns)
        last_paragraph = body.findall('.//text:p', ns)[-1]

        # 遍历剩余ODT文件,提取并合并内容
        for odt_file in input_files[1:]:
            with zipfile.ZipFile(odt_file, 'r') as zip_ref:
                # 临时提取当前文件的content.xml
                temp_dir = os.path.join(tmpdir, 'temp')
                os.mkdir(temp_dir)
                zip_ref.extract('content.xml', temp_dir)
                temp_content = os.path.join(temp_dir, 'content.xml')
                temp_tree = etree.parse(temp_content)
                temp_root = temp_tree.getroot()
                temp_body = temp_root.find('.//text:body', ns)

                # 提取正文内的段落、表格等元素,排除末尾空段落避免多余空白
                for elem in temp_body.findall('.//text:p | .//text:table', ns)[:-1]:
                    body.insert(body.index(last_paragraph) + 1, elem)
            
            # 清理临时文件
            os.remove(temp_content)
            os.rmdir(temp_dir)
        
        # 重新打包为ODT文件
        with zipfile.ZipFile(output_file, 'w', zipfile.ZIP_DEFLATED) as output_zip:
            for foldername, subfolders, filenames in os.walk(tmpdir):
                for filename in filenames:
                    file_path = os.path.join(foldername, filename)
                    arcname = os.path.relpath(file_path, tmpdir)
                    output_zip.write(file_path, arcname)

# 使用示例
input_odts = ['report_test_1_fulled.odt', 'report_test_2_fulled.odt']
merge_odts(input_odts, 'merged_output.odt')

方案二:转DOCX后合并(适配DOCX输出需求)

如果目标是生成DOCX,可通过LibreOffice命令行将ODT转成DOCX,再用python-docx合并:

  1. 安装依赖:

    • 安装LibreOffice(macOS用Homebrew):
      brew install libreoffice
      
    • 安装python-docx:
      pip install python-docx
      
  2. 转换+合并代码示例:

import os
import subprocess
from docx import Document

def odt_to_docx(odt_path):
    docx_path = os.path.splitext(odt_path)[0] + '.docx'
    # 调用LibreOffice无头模式转换文件
    subprocess.run([
        '/Applications/LibreOffice.app/Contents/MacOS/soffice',
        '--headless', '--convert-to', 'docx', odt_path,
        '--outdir', os.path.dirname(odt_path)
    ], check=True)
    return docx_path

def merge_docxs(input_docxs, output_docx):
    merged_doc = Document()
    # 复制第一个文档的样式到合并文档
    first_doc = Document(input_docxs[0])
    for style in first_doc.styles:
        if style not in merged_doc.styles:
            merged_doc.styles.add(style)
    
    # 逐个合并文档内容
    for idx, docx_path in enumerate(input_docxs):
        doc = Document(docx_path)
        if idx != 0:
            # 插入分页符分隔不同源文件内容
            merged_doc.add_page_break()
        # 复制当前文档的所有正文元素
        for element in doc.element.body:
            merged_doc.element.body.append(element)
    
    merged_doc.save(output_docx)

# 使用示例
input_odts = ['report_test_1_fulled.odt', 'report_test_2_fulled.odt']
docx_files = [odt_to_docx(odt) for odt in input_odts]
merge_docxs(docx_files, 'merged_output.docx')

# 可选:删除转换生成的临时DOCX文件
for docx in docx_files:
    os.remove(docx)

关于Aspose安装失败的补充

你遇到的aspose-words安装失败,可尝试指定兼容macOS Catalina的版本:

pip install aspose-words==23.10.0

若仍无法安装,优先选择上述两个免费方案,无需依赖付费商业库。


内容的提问来源于stack exchange,提问作者MAP

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 23:45:42