You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python将docx/doc(ioBytes格式)转换为PDF(ioBytes格式)

解决方案

原代码问题分析

  • 错误创建空文档覆盖原文件内容:你用Document()生成了空文档,再将其保存到in_file,直接覆盖了之前从本地读取的docx内容,正确做法应该是用Document(in_file)加载传入的BytesIO文件。
  • docx2pdf不支持直接操作内存对象:该库底层依赖系统的Word(Windows)或LibreOffice(Linux/macOS)完成转换,这些工具只支持实体文件路径,无法直接处理Document对象和BytesIO输出。

核心解决思路

通过临时文件中转处理内存中的BytesIO数据:

  1. 将云存储读取到的BytesIO内容写入临时docx文件
  2. 调用转换工具将临时docx转成临时PDF文件
  3. 读取临时PDF文件到BytesIO,再上传回云存储
  4. 清理临时文件

兼容多系统的代码实现

依赖安装

  • 安装Python库:
    pip install python-docx docx2pdf tempfile
    
  • 系统依赖:
    • Windows:需安装Microsoft Word
    • Linux:执行sudo apt install libreoffice安装LibreOffice
    • macOS:用brew执行brew install libreoffice安装LibreOffice

修正代码

import io
import tempfile
import os
from docx import Document
from docx2pdf import convert

# 模拟从云存储读取文件到BytesIO(实际替换为你的云存储读取逻辑)
with open('prueba.docx', 'rb') as f:
    in_file = io.BytesIO(f.read())

# 将内存中的docx写入临时文件
with tempfile.NamedTemporaryFile(suffix='.docx', delete=False) as temp_docx:
    temp_docx.write(in_file.getvalue())
    temp_docx_path = temp_docx.name

try:
    # 转换docx到PDF临时文件
    temp_pdf_path = temp_docx_path.replace('.docx', '.pdf')
    convert(temp_docx_path, temp_pdf_path)

    # 读取PDF到BytesIO(准备上传云存储)
    output = io.BytesIO()
    with open(temp_pdf_path, 'rb') as f:
        output.write(f.read())
    output.seek(0)

    # 此处添加你的云存储上传逻辑,例如:
    # cloud_storage_client.upload_blob(output, '目标存储路径/prueba.pdf')

finally:
    # 清理临时文件
    os.unlink(temp_docx_path)
    if os.path.exists(temp_pdf_path):
        os.unlink(temp_pdf_path)

纯Python依赖替代方案(无系统Office)

如果不想依赖系统级Office软件,可使用python-docx结合pdfkit,但对复杂格式(图片、样式、页眉页脚)支持有限,适合简单文档:

  1. 安装依赖:

    pip install python-docx pdfkit
    
  2. 安装系统依赖wkhtmltopdf:

    • Windows:下载安装后添加到环境变量
    • Ubuntu:sudo apt install wkhtmltopdf
    • macOS:brew install wkhtmltopdf
  3. 代码示例

import io
from docx import Document
import pdfkit

# 读取云存储文件到BytesIO
with open('prueba.docx', 'rb') as f:
    in_file = io.BytesIO(f.read())

doc = Document(in_file)

# 将docx内容转为HTML(pdfkit支持HTML转PDF)
html_content = '<html><body>'
for para in doc.paragraphs:
    html_content += f'<p>{para.text}</p>'
for table in doc.tables:
    html_content += '<table border="1">'
    for row in table.rows:
        html_content += '<tr>'
        for cell in row.cells:
            html_content += f'<td>{cell.text}</td>'
        html_content += '</tr>'
    html_content += '</table>'
html_content += '</body></html>'

# 转换HTML到PDF BytesIO
output = io.BytesIO()
pdfkit.from_string(html_content, output)
output.seek(0)

# 上传到云存储
# cloud_storage_client.upload_blob(output, '目标存储路径/prueba.pdf')

内容的提问来源于stack exchange,提问作者Daniela

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 04:01:12