Python实现邮件PDF附件下载及与邮件内容PDF合并的技术问题求助
Fixing PDF Attachment Saving & Merging in Your Email Processing Script
Hey there! I see you're stuck on saving PDF attachments from emails and merging them with the PDF generated from the email body. Let's get that sorted out.
First, let's break down the gaps in your current code:
- The
fileslist gets reinitialized inside thefor part in email_message_raw.walk()loop, so it'll only keep the last file you add instead of collecting all relevant PDFs. - You're missing the core logic to write the PDF attachment's binary content to a file in the
pdfcontent_type block.
Here's the revised code with fixes and clear explanations:
import imaplib import email from PyPDF2 import PdfFileMerger import os from uuid import uuid4 # Generate unique filenames to avoid overwriting existing files HOST = ***** USERNAME = **** PASSWORD = ***** m = imaplib.IMAP4_SSL(HOST, 993) m.login(USERNAME, PASSWORD) m.select('INBOX') result, data = m.uid('search', None, "ALL") if result == 'OK': for num in data[0].split(): result, data = m.uid('fetch', num, '(RFC822)') if result == 'OK': email_message_raw = email.message_from_bytes(data[0][1]) files = [] # Move this outside the part loop to persist across all email parts email_uid = num.decode('utf-8') # Grab email UID for unique filename tracking for part in email_message_raw.walk(): content_type = part.get_content_type() if "plain" in content_type: text = part.get_payload() # Write plain text to HTML and convert to PDF with open('text.html','w') as f: f.write('<p>'+text+'</p>') os.system('wkhtmltopdf text.html text.pdf') os.remove('text.html') # Safer single-file deletion than rm -r files.append('text.pdf') if "pdf" in content_type: # Get the original attachment filename filename = part.get_filename() # Generate a unique name if the attachment has no filename if not filename: filename = f"email_{email_uid}_attachment_{uuid4().hex}.pdf" # Save the decoded PDF attachment to disk with open(filename, 'wb') as f: f.write(part.get_payload(decode=True)) # Decode encoded attachment content files.append(filename) print(f"Saved PDF attachment: {filename}") # Merge PDFs only if we have files to combine if files: merger = PdfFileMerger() for pdf in files: merger.append(pdf) # Name merged file with email UID to avoid overwrites merger.write(f'complete_email_{email_uid}.pdf') merger.close() # Clean up temporary files after merging for pdf in files: if os.path.exists(pdf): os.remove(pdf) else: print(f"No PDFs found to merge for email UID {email_uid}") m.close() m.logout()
Key Improvements Explained:
- Persistent file collection: Moving
files = []outside the part loop ensures we gather all PDFs (body-generated and attachments) instead of resetting the list every time we process a new email part. - PDF attachment saving:
- Uses
part.get_filename()to retain the original attachment name when available. - Generates a unique fallback filename using the email UID and UUID to prevent overwriting existing files.
- Decodes the attachment payload with
part.get_payload(decode=True)—critical because email attachments are encoded in base64 or similar formats.
- Uses
- Cleaner file management:
- Replaced
os.system('rm -r')withos.remove()for safer, more reliable single-file deletion. - Automatically cleans up temporary PDFs after merging to avoid cluttering your directory.
- Replaced
- Unique merged filenames: Appends the email UID to the merged PDF name so you don't accidentally overwrite merged files from different emails.
This should resolve your attachment saving issue and make the merging workflow work as expected. Let me know if you run into any other hiccups!
内容的提问来源于stack exchange,提问作者tedioustortoise
相关产品推荐
相关产品推荐

