Python添加PDF水印报错TypeError:传入PdfFileReader而非路径对象
Let's break down what's going wrong here and fix it step by step:
The Root Cause
Your add_watermark function is written to expect a file path string (like 'watermark.pdf') as its first argument—it tries to open that path with open(wmFile, 'rb'). But in your main function, you're passing mywatermark, which is already a PyPDF2.PdfFileReader object (created from an in-memory BytesIO buffer). That's why you get the TypeError: open() can't work with a PdfFileReader, only path-like objects.
The Fixed Code
We'll rewrite the add_watermark function to accept a pre-made PdfFileReader instead of a file path, since you're generating your watermark in memory. Here's the corrected full code:
import PyPDF2 import io from reportlab.pdfgen import canvas from reportlab.lib.pagesizes import letter def add_watermark(wm_pdf_reader, pageObj): # Directly use the passed PdfFileReader object instead of opening a file wm_page = wm_pdf_reader.getPage(0) pageObj.mergePage(wm_page) return pageObj def main(): # Create in-memory watermark PDF with ReportLab packet = io.BytesIO() can = canvas.Canvas(packet, pagesize=letter) can.setFont('Helvetica-Bold', 18) # Replace this static text with CSV content later can.drawString(10, 100, "Hello world") can.showPage() can.save() # Reset buffer to start and create PdfFileReader packet.seek(0) mywatermark = PyPDF2.PdfFileReader(packet) origFileName = 'Module1.pdf' newFileName = 'watermarked_example.pdf' # Process original PDF with open(origFileName, 'rb') as pdfFileObj: pdfReader = PyPDF2.PdfFileReader(pdfFileObj) pdfWriter = PyPDF2.PdfFileWriter() # Add watermark to each page for page_num in range(pdfReader.numPages): page = pdfReader.getPage(page_num) wmpageObj = add_watermark(mywatermark, page) pdfWriter.addPage(wmpageObj) # Save the output PDF with open(newFileName, 'wb') as newFile: pdfWriter.write(newFile) if __name__ == "__main__": main()
Key Changes Made
- Updated
add_watermark: Removed file opening logic—now it takes aPdfFileReaderdirectly, grabs its first page, and merges it with the target page. - Used context managers (
withstatements): Safer file handling that automatically closes files, no need for manualclose()calls. - Cleaned up imports: Moved top-level imports out of
main()to avoid redundant imports.
Prepping for CSV Integration
When you're ready to replace "Hello world" with CSV content:
- Use Python's
csvmodule to read your CSV file (e.g.,csv.readerorcsv.DictReader). - If you want different text per page, map CSV rows to PDF pages (make sure your CSV has enough rows for all pages).
- Modify the
drawStringline to use the CSV data instead of the static string. For example:# Example: Read CSV and use first column for page watermarks import csv with open('watermark_texts.csv', 'r') as csv_file: csv_reader = csv.reader(csv_file) watermark_texts = [row[0] for row in csv_reader] # Then in the canvas section, loop per page (adjust as needed) # For a single text across all pages, just pick one entry from watermark_texts
内容的提问来源于stack exchange,提问作者f0rd42

