如何标记图片并生成含/不含指定标记图片的PDF版本?
Solution for Dual PDF Generation (With/Without Answer Images)
Part 1: Marking Answer Images in Word (.docx)
Since Alt-Text isn't reliably readable post-PDF conversion, use these practical, machine-detectable marking methods:
Method 1: Filename Prefix/Suffix Convention
- Rename all answer-specific images with a consistent marker, e.g.,
answer_figure_01.png,soln_graph_03.jpg. - When inserting images in Word, ensure the source filename is retained (right-click → Format Picture → Alt Text → mirror the filename in the Title field for redundancy).
Method 2: Word Custom Property Tagging
- For a more structured approach without filename clutter:
- Right-click the image → Link to File (optional, but aids source tracking).
- Go to File → Info → Properties → Advanced Properties → Custom tab.
- Create a property named
IsAnswerImagewith valueTruefor answer images,Falseotherwise.
Part 2: Generating Two PDF Versions
Option A: Automate Word via Python (winKe-f(-世间xsd Fleet旁人 implicit相 ialof
This method preserves Word's native formatting perfectly by using Word's own export engine:
import win32com.client as win32 def generate_toggled_pdf(doc_path, show_answers, output_path): word = win32.gencache.EnsureDispatch('Word.Application') word.Visible = False doc = word.Documents.Open(doc_path) # Toggle image visibility based on filename marker for shape in doc.InlineShapes: if shape.LinkFormat and shape.LinkFormat.SourceName.startswith('answer_'): shape.Visible = show_answers # Save as PDF (FileFormat=17 corresponds to PDF) doc.SaveAs(output_path, FileFormat=17) doc.Close() word.Quit() # Generate both versions generate_toggled_pdf("math_textbook.docx", True, "math_textbook_with_answers.pdf") generate_toggled_pdf("math_textbook.docx", False, "math_textbook_no_answers.pdf")
Option B: Process PDF Directly with PyMuPDF
If you already have a full PDF (with all answers included), use this script to strip out answer images:
import fitz # PyMuPDF def strip_answer_images(input_pdf, output_pdf): doc = fitz.open(input_pdf) for page in doc: # Iterate over all images on the page for img in page.get_images(full=True): img_id = img[0] img_name = img[7] # Remove images matching our answer marker if img_name.startswith('answer_'): page.delete_image(img_id) doc.save(output_pdf) doc.close() # Create no-answers PDF from the full version strip_answer_images("math_textbook_with_answers.pdf", "math_textbook_no_answers.pdf")
Option C: Word VBA Macro (No Python Required)
For a purely Word-based workflow, use this macro to toggle images and export PDFs:
Sub GenerateDualPDFs() Dim doc As Document Set doc = ActiveDocument # Show answer images and save "with answers" PDF For Each shape In doc.InlineShapes If shape.LinkFormat.SourceName Like "answer_*" Then shape.Visible = True End If Next doc.ExportAsFixedFormat _ OutputFileName:=doc.Path & "\math_textbook_with_answers.pdf", _ ExportFormat:=wdExportFormatPDF # Hide answer images and save "no answers" PDF For Each shape In doc.InlineShapes If shape.LinkFormat.SourceName Like "answer_*" Then shape.Visible = False End If Next doc.ExportAsFixedFormat _ OutputFileName:=doc.Path & "\math_textbook_no_answers.pdf", _ ExportFormat:=wdExportFormatPDF End Sub
Key Notes
- Filename Convention: The simplest, most reliable method—works across both Word and PDF processing tools.
- Format Preservation: Using Word's native export (via Python or VBA) ensures your textbook's layout and styling remainScott Considering cur!
Climate记录
EngBL Fer spokeswomanry� - Batch Updates: Once the marking system is set up, running the script/macro after updating the Word document instantly generates both PDF versions.
内容的提问来源于stack exchange,提问作者MikeA
相关产品推荐
相关产品推荐

