You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何扫描发票提取指定姓名信息导入Word模板,优先采用Python实现

基于Python的发票姓名提取+Word批量打印解决方案

核心实现逻辑

整个流程仅需要用户做1次操作:导入所有扫描完成的发票文件,后续步骤全自动执行。

  • 步骤1:自动识别所有发票文件中的姓名信息
  • 步骤2:将提取到的姓名批量填入预设的Word模板,生成独立的Word文件
  • 步骤3:自动调用系统默认Office程序打开所有生成的Word文件,用户直接点击打印即可

依赖库安装

执行以下命令安装所需依赖:

pip install pdfplumber pytesseract pillow docxtpl python-docx

如果是扫描件/OCR场景,需要提前安装Tesseract OCR引擎,Windows平台直接下载安装包配置环境变量即可,macOS/Linux可通过brew/apt直接安装。

分模块实现代码

1. 发票姓名提取模块

支持电子发票PDF和扫描件/图片格式发票:

import re
import pdfplumber
from PIL import Image
import pytesseract

# 匹配姓名的正则,可根据自己的发票格式调整规则
NAME_PATTERN = re.compile(r"(购买方\s*[::]\s*)([\u4e00-\u9fa5]{2,4})|(收货人\s*[::]\s*)([\u4e00-\u9fa5]{2,4})")

def extract_name_from_invoice(file_path):
    # 处理PDF格式电子发票
    if file_path.endswith(".pdf"):
        with pdfplumber.open(file_path) as pdf:
            text = "\n".join([page.extract_text() for page in pdf.pages if page.extract_text()])
    # 处理图片格式扫描发票
    elif file_path.endswith((".png", ".jpg", ".jpeg")):
        img = Image.open(file_path)
        text = pytesseract.image_to_string(img, lang="chi_sim")
    else:
        return None
    # 正则匹配提取姓名
    match_res = NAME_PATTERN.search(text)
    if match_res:
        return match_res.group(2) or match_res.group(4)
    return None

2. Word模板批量填充模块

提前在Word模板中需要插入姓名的位置标注{{user_name}},保存为template.docx:

from docxtpl import DocxTemplate
import os

OUTPUT_FOLDER = "./generated_docs"
os.makedirs(OUTPUT_FOLDER, exist_ok=True)

def generate_word_from_template(name):
    doc = DocxTemplate("template.docx")
    context = {"user_name": name}
    doc.render(context)
    output_path = os.path.join(OUTPUT_FOLDER, f"{name}_打印文件.docx")
    doc.save(output_path)
    return output_path

3. 自动打开生成的Word文件

import platform

def open_file(file_path):
    if platform.system() == "Windows":
        os.startfile(file_path)
    elif platform.system() == "Darwin":  # macOS
        os.system(f"open {file_path}")
    else:  # Linux
        os.system(f"xdg-open {file_path}")

主流程调用

if __name__ == "__main__":
    # 扫描存放发票的文件夹,可根据自己的路径调整
    INVOICE_FOLDER = "./invoices"
    all_names = []
    # 批量提取姓名
    for file_name in os.listdir(INVOICE_FOLDER):
        file_path = os.path.join(INVOICE_FOLDER, file_name)
        name = extract_name_from_invoice(file_path)
        if name:
            all_names.append(name)
    # 批量生成Word
    generated_files = []
    for name in all_names:
        file_path = generate_word_from_template(name)
        generated_files.append(file_path)
    # 一次性打开所有生成的Word文件
    for file in generated_files:
        open_file(file)

优化建议

如果怕OCR识别错误,可以在提取完所有姓名后加个简单的交互校验,比如用tkinter做个极简列表展示提取到的姓名,用户确认无误后再生成Word,不会增加太多操作成本。

内容的提问来源于stack exchange,提问作者Shaz_

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 03:45:04