PyMuPDF处理PDF插入图像质量下降问题排查及代码求助
PDF页面添加空白区域后画质下降的原因
我要给PDF每页的顶部/底部加空白区域,做法是提取页面图像,用Pillow生成调整高度的新图,删掉原页图像再插入新图,但做完后图像质量明显变差。下面是原始代码和优化后的代码,想问下问题出在哪?
原始代码
import fitz # PyMuPDF from PIL import Image from io import BytesIO def flatten_images_in_pdf(byte_array, placementType, verticalPosition, text_height): # Open the PDF from the byte array pdf_document = fitz.open(stream=byte_array, filetype="pdf") # loop through pages for page_number in range(len(pdf_document)): # Get the page page = pdf_document[page_number] page_pixmap = page.get_pixmap() # Determine the color mode based on the number of components if page_pixmap.n == 1: color_mode = 'L' # Grayscale or black and white elif page_pixmap.n == 3: color_mode = 'RGB' # Color else: color_mode = 'CMYK' # CMYK or other color modes page_pil_image = Image.frombytes(color_mode, [int(page_pixmap.width), int(page_pixmap.height)], page.get_pixmap().samples) # Get the dimensions (width and height) of the image in pixels width_pixels, height_pixels = page_pil_image.size # Get the dimensions of the media box in points media_box = page.mediabox width_points = media_box[2] height_points = media_box[3] # Calculate the DPI dpi_x = width_pixels / (width_points / 72) # 72 points = 1 inch dpi_y = height_pixels / (height_points / 72) page_edge_offset = 0.5 if (placementType.lower() == "margin"): # The margin will be the page edge offset and the height of the stamp margin_height = (page_edge_offset * dpi_y) + text_height; new_height = int(page_pixmap.height - margin_height); # Check for invalid new height if (new_height<= 0): raise Exception ("New height for page is less than 0") # Create a new blank image with the adjusted height new_img = Image.new(color_mode, (page_pixmap.width, page_pixmap.height), white) # Determine the position to paste the old image onto the new one if verticalPosition == "top": # if the text position is top, the image needs to start from the bottom position = (0, 0) else: # if the text position is bottom, the image needs to start from the top position = (0, new_height - height_pixels) # Paste the old image onto the new one new_img.paste(page_pil_image, position) # Convert the modified Pillow image to a bytes-like object (e.g., PNG format) image_bytes = BytesIO() new_img.save(image_bytes, format="GIF",dpi=(dpi_x,dpi_y)) image_bytes.seek(0) images_in_page = page.get_images() for image in images_in_page: image_xref = image[0] # the xref is the first property. page.delete_image(image_xref) page.insert_image(rect=page.rect, stream = image_bytes) # Create an in-memory byte stream output_stream = BytesIO() # Save the modified PDF to the byte stream pdf_document.save(output_stream) pdf_document.close()
优化后代码(移除手动色彩模式判断,改用Pixmap转PPM后用Pillow打开)
# Open the PDF from the byte array pdf_document = fitz.open(stream=byte_array, filetype="pdf") # loop through pages for page_number in range(len(pdf_document)): # Get the page page = pdf_document[page_number] bate_stamp = bate_stamps[page_number] page_pixmap = page.get_pixmap() page_pil_image = Image.open(BytesIO(page_pixmap.tobytes("ppm"))) # Get the dimensions of the media box in points width_points = page.mediabox[2] height_points = page.mediabox[3] # Calculate the DPI dpi_x = page_pil_image.size[0] / (width_points / 72) # 72 points = 1 inch dpi_y = page_pil_image.size[1] / (height_points / 72) page_edge_offset = 0.5 if (placementType.lower() == "margin"): # The margin will be the page edge offset and the height of the stamp margin_height = (page_edge_offset * dpi_y) + text_height; new_height = int(page_pixmap.height - margin_height); # Check for invalid new height if (new_height<= 0): raise Exception ("New height for page when setting the bate stam as margin is less than 0") # Create a new blank image with the adjusted height new_img = Image.new(page_pil_image.mode, (page_pixmap.width, page_pixmap.height), white) # Determine the position to paste the old image onto the new one if verticalPosition == "top": # if the batestamp position is top, the image needs to start from the bottom position = (0, 0) else: # if the batestamp position is bottom, the image needs to start from the top position = (0, new_height - page_pil_image.size[1]) # Paste the old image onto the new one new_img.paste(page_pil_image, position) # Convert the modified Pillow image to a bytes-like object (e.g., PNG format) image_bytes = BytesIO() new_img.save(image_bytes, format="GIF",dpi=(dpi_x,dpi_y)) image_bytes.seek(0) images_in_page = page.get_images() for image in images_in_page: image_xref = image[0] # the xref is the first property. page.delete_image(image_xref) page.insert_image(rect=page.rect, stream = image_bytes) # Create an in-memory byte stream output_stream = BytesIO() # Save the modified PDF to the byte stream pdf_document.save(output_stream) pdf_document.close()
问题原因解析
1. GIF格式是画质损失的核心元凶
两代代码都用format="GIF"保存图像,但GIF本身是索引色模式,最多只能显示256种颜色,远低于PDF页面原生的RGB/CMYK全色域。这会直接导致:
- 色彩丰富的页面出现严重色偏、色阶断裂
- 渐变、阴影等精细效果丢失细节
- 不支持半透明通道(仅1位透明),透明区域处理失真
2. 原始代码的色彩模式匹配错误
原始代码手动通过page_pixmap.n判断色彩模式:
- 当
page_pixmap.n == 4时(比如带Alpha通道的RGBA图像),会被错误归类为CMYK模式,Image.frombytes解析时通道顺序完全错位,直接导致图像模糊、偏色 - PyMuPixmap的
samples数据格式和Pillowfrombytes的预期格式不完全匹配,手动转换容易出现字节对齐错误,进一步降低画质
3. 优化代码仅解决了部分问题
优化代码改用page_pixmap.tobytes("ppm")让Pillow自动识别色彩模式,解决了通道错位的问题,但仍然使用GIF格式保存,所以色彩压缩导致的画质损失依然存在。
内容的提问来源于stack exchange,提问作者Style
相关产品推荐
相关产品推荐

