如何使用Python将PPT转PDF时正确提取并嵌入图片?
PPT转PDF的Python Flask函数实现
以下代码是一个Python函数,接收PowerPoint文件(ppt_file),借助多个第三方库和Python内置模块将其转换为PDF文件。以下是该代码的分步执行说明:
- 使用
tempfile模块创建带.pptx扩展名的临时文件,将上传的ppt_file保存到该临时文件中。 - 再次使用
tempfile模块创建带.pdf扩展名的临时文件。 - 使用
python-pptx库读取第一步创建的临时文件中的PowerPoint文件。 - 使用
reportlab库创建新的PDF输出文件。 - 针对PowerPoint中的每张幻灯片,计算幻灯片尺寸,并使用
reportlab的landscape()或portrait()函数设置PDF画布大小。 - 将幻灯片内容提取为HTML,再使用
reportlab函数绘制到PDF画布上。 - 如果幻灯片包含图片,使用
Pillow库将图片转换为PNG文件并绘制到PDF画布上。 - 处理完所有幻灯片后,将PDF文件保存到第二步创建的临时文件中。
- 将临时PDF文件读取到
BytesIO对象中,通过send_file以Flask响应形式返回。 - 最后使用
os.unlink删除临时文件。
该函数如下:
import io import os import tempfile from pptx import Presentation from reportlab.lib.pagesizes import landscape, portrait from reportlab.pdfgen import canvas from flask import send_file import imgkit from PIL import Image def ppt_to_pdf(ppt_file): # Save the uploaded file to a temporary file with tempfile.NamedTemporaryFile(delete=False, suffix=".pptx") as temp_ppt_file: ppt_file.save(temp_ppt_file.name) temp_ppt_file.close() # Create a temporary PDF file with tempfile.NamedTemporaryFile(delete=False, suffix=".pdf") as temp_pdf_file: temp_pdf_file.close() # Read the PowerPoint file prs = Presentation(temp_ppt_file.name) # Create a PDF file using reportlab pdf = canvas.Canvas(temp_pdf_file.name) for slide in prs.slides: slide_width = prs.slide_width.pt slide_height = prs.slide_height.pt # Set the page size based on the slide dimensions if slide_width > slide_height: pdf.setPageSize(landscape((slide_width, slide_height))) else: pdf.setPageSize(portrait((slide_width, slide_height))) # Extract slide content as HTML slide_html = "<html><body>" for shape in slide.shapes: if shape.has_text_frame: slide_html += f"<p>{shape.text}</p>" elif shape.shape_type == 13: # Picture with tempfile.NamedTemporaryFile(delete=False, suffix=".png") as temp_image_file: temp_image_file.write(shape.image.blob) temp_image_file.close() img = Image.open(temp_image_file.name) # Calculate the position and size of the image on the slide left = shape.left.pt top = shape.top.pt width = shape.width.pt height = shape.height.pt # Draw the image on the PDF canvas pdf.drawImage(temp_image_file.name, left, top, width=width, height=height) # Clean up the temporary image file os.unlink(temp_image_file.name) slide_html += "</body></html>" # Convert slide content to an image using imgkit slide_image = imgkit.from_string( slide_html, False, {"width": int(slide_width), "height": int(slide_height)}) # Save the slide image to a temporary file with tempfile.NamedTemporaryFile(delete=False, suffix=".png") as temp_image_file: temp_image_file.write(slide_image) temp_image_file.close() # Draw the slide image on the PDF canvas pdf.drawImage(temp_image_file.name, 0, 0, width=slide_width, height=slide_height) # Clean up the temporary image file os.unlink(temp_image_file.name) # Move to the next page pdf.showPage() pdf.save() # Read the temporary PDF file into a BytesIO object final_pdf = io.BytesIO() with open(temp_pdf_file.name, 'rb') as f: final_pdf.write(f.read()) final_pdf.seek(0) # Clean up the temporary files os.unlink(temp_ppt_file.name) os.unlink(temp_pdf_file.name) # Return the PDF file as a Flask response return send_file(final_pdf, download_name='output.pdf', as_attachment=True, mimetype='application/pdf')
期望该函数能返回PPT文件对应的PDF版本。
内容的提问来源于stack exchange,提问作者falfalaye GPT
相关产品推荐
相关产品推荐

