批量PDF发货单处理:制作带背面条款模板并支持批处理调用打印
Let’s break this down into practical, batch-friendly steps using accessible tools:
1. Build the Invoice Template (with Backside Terms)
First, create a reusable template that pairs your invoice layout on the front with terms/conditions on the back:
- Use LibreOffice Writer or Microsoft Word to design the front page: Add clear placeholders like
{{CustomerName}},{{OrderNumber}},{{ShippingAddress}}(these will get replaced with extracted data later). - For the back page, paste your full terms and conditions—format it for easy readability when printed.
- Save the template as a DOCX (for mail merge workflows) or a PDF form (use LibreOffice’s "Export as PDF" with "Create PDF form" enabled, or Adobe Acrobat to add named fields).
If you prefer code-based control, use Jinja2 (Python templating) to make an HTML template with CSS page breaks separating front and back pages.
2. Extract Data from Source PDFs
Turn your raw PDF data into a structured format (CSV/JSON) that’s easy to merge. Here are two reliable methods:
Option A: Command-Line Parsing (for consistent PDFs)
Use pdftotext (part of Poppler tools) to extract text, then parse it into CSV with awk:
# Extract text from one PDF pdftotext source_invoice.pdf temp.txt # Parse into CSV (adjust regex to match your PDF's layout) awk '/Customer Name:/ {name=$3" "$4} /Order Number:/ {order=$3} /Shipping Address:/ {addr=$3" "$4" "$5} END {print name","order","addr}' temp.txt >> extracted_data.csv
Loop through all source PDFs in a shell/batch script to build a full CSV of invoice data.
Option B: Python Script (for complex/unstructured PDFs)
Use pdfplumber to extract structured data with more flexibility:
import pdfplumber import csv import os def extract_invoice_data(pdf_path): data = {} with pdfplumber.open(pdf_path) as pdf: page = pdf.pages[0] text = page.extract_text() # Pull fields using string matching for line in text.split('\n'): if 'Customer Name:' in line: data['CustomerName'] = line.split(':')[1].strip() if 'Order Number:' in line: data['OrderNumber'] = line.split(':')[1].strip() # Add more fields as needed return data # Process all PDFs in a folder with open('extracted_data.csv', 'w', newline='') as csvfile: fieldnames = ['CustomerName', 'OrderNumber', 'ShippingAddress'] writer = csv.DictWriter(csvfile, fieldnames=fieldnames) writer.writeheader() for filename in os.listdir('source_pdfs'): if filename.endswith('.pdf'): data = extract_invoice_data(os.path.join('source_pdfs', filename)) writer.writerow(data)
3. Merge Data into the Template (Batch-Friendly)
Choose a method that works via command line for automated processing:
Option A: LibreOffice Command-Line Mail Merge
If you have a DOCX template with mail merge fields:
# Generate individual PDFs for each row in your CSV libreoffice --headless --convert-to pdf --outdir output_invoices "shipping_template.docx" --merge extracted_data.csv
This creates separate invoices with front layout and back terms for every entry.
Option B: Python + Jinja2 + WeasyPrint
For full control over styling:
- Create a Jinja2 HTML template (
invoice_template.html) with CSS page breaks for the backside. - Use WeasyPrint to convert filled templates to PDF:
from jinja2 import Environment, FileSystemLoader from weasyprint import HTML import csv env = Environment(loader=FileSystemLoader('.')) template = env.get_template('invoice_template.html') with open('extracted_data.csv', 'r') as csvfile: reader = csv.DictReader(csvfile) for row in reader: html_output = template.render(row) HTML(string=html_output).write_pdf(f"output_invoices/invoice_{row['OrderNumber']}.pdf")
Option C: PDFtk (Fill PDF Form Fields)
If you made a PDF form with named fields:
# Fill a single form pdftk shipping_template.pdf fill_form extracted_data.csv output output_invoices/invoice.pdf flatten # Batch process (Windows batch script example) @echo off for /f "tokens=*" %%a in ('dir /b source_pdfs\*.pdf') do ( pdftk shipping_template.pdf fill_form extracted_data.csv output output_invoices\%%~na_invoice.pdf flatten )
4. Automate with Batch Processing
Wrap all steps into a single script for one-click execution:
- Windows: Create a
.batfile that runs extraction, merging, and cleanup. - Linux/macOS: Use a
.shshell script.
Example Windows batch script:
@echo off :: Step 1: Extract data from source PDFs python extract_data.py :: Step 2: Merge data into template libreoffice --headless --convert-to pdf --outdir output_invoices "shipping_template.docx" --merge extracted_data.csv echo Batch processing complete! Check the output_invoices folder.
Schedule this script with Windows Task Scheduler or cron (Linux/macOS) for fully automated runs.
内容的提问来源于stack exchange,提问作者mary

