如何用Python将docx/doc(ioBytes格式)转换为PDF(ioBytes格式)
解决方案
原代码问题分析
- 错误创建空文档覆盖原文件内容:你用
Document()生成了空文档,再将其保存到in_file,直接覆盖了之前从本地读取的docx内容,正确做法应该是用Document(in_file)加载传入的BytesIO文件。 docx2pdf不支持直接操作内存对象:该库底层依赖系统的Word(Windows)或LibreOffice(Linux/macOS)完成转换,这些工具只支持实体文件路径,无法直接处理Document对象和BytesIO输出。
核心解决思路
通过临时文件中转处理内存中的BytesIO数据:
- 将云存储读取到的BytesIO内容写入临时docx文件
- 调用转换工具将临时docx转成临时PDF文件
- 读取临时PDF文件到BytesIO,再上传回云存储
- 清理临时文件
兼容多系统的代码实现
依赖安装
- 安装Python库:
pip install python-docx docx2pdf tempfile - 系统依赖:
- Windows:需安装Microsoft Word
- Linux:执行
sudo apt install libreoffice安装LibreOffice - macOS:用brew执行
brew install libreoffice安装LibreOffice
修正代码
import io import tempfile import os from docx import Document from docx2pdf import convert # 模拟从云存储读取文件到BytesIO(实际替换为你的云存储读取逻辑) with open('prueba.docx', 'rb') as f: in_file = io.BytesIO(f.read()) # 将内存中的docx写入临时文件 with tempfile.NamedTemporaryFile(suffix='.docx', delete=False) as temp_docx: temp_docx.write(in_file.getvalue()) temp_docx_path = temp_docx.name try: # 转换docx到PDF临时文件 temp_pdf_path = temp_docx_path.replace('.docx', '.pdf') convert(temp_docx_path, temp_pdf_path) # 读取PDF到BytesIO(准备上传云存储) output = io.BytesIO() with open(temp_pdf_path, 'rb') as f: output.write(f.read()) output.seek(0) # 此处添加你的云存储上传逻辑,例如: # cloud_storage_client.upload_blob(output, '目标存储路径/prueba.pdf') finally: # 清理临时文件 os.unlink(temp_docx_path) if os.path.exists(temp_pdf_path): os.unlink(temp_pdf_path)
纯Python依赖替代方案(无系统Office)
如果不想依赖系统级Office软件,可使用python-docx结合pdfkit,但对复杂格式(图片、样式、页眉页脚)支持有限,适合简单文档:
安装依赖:
pip install python-docx pdfkit安装系统依赖wkhtmltopdf:
- Windows:下载安装后添加到环境变量
- Ubuntu:
sudo apt install wkhtmltopdf - macOS:
brew install wkhtmltopdf
代码示例
import io from docx import Document import pdfkit # 读取云存储文件到BytesIO with open('prueba.docx', 'rb') as f: in_file = io.BytesIO(f.read()) doc = Document(in_file) # 将docx内容转为HTML(pdfkit支持HTML转PDF) html_content = '<html><body>' for para in doc.paragraphs: html_content += f'<p>{para.text}</p>' for table in doc.tables: html_content += '<table border="1">' for row in table.rows: html_content += '<tr>' for cell in row.cells: html_content += f'<td>{cell.text}</td>' html_content += '</tr>' html_content += '</table>' html_content += '</body></html>' # 转换HTML到PDF BytesIO output = io.BytesIO() pdfkit.from_string(html_content, output) output.seek(0) # 上传到云存储 # cloud_storage_client.upload_blob(output, '目标存储路径/prueba.pdf')
内容的提问来源于stack exchange,提问作者Daniela
相关产品推荐
相关产品推荐

