You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python从批量PDF提取固定位置员工ID并重命名文件的技术咨询

问题与解决思路

问题描述

我用这段代码批量提取文件夹中PDF的文本:

from pypdf import PdfReader
import os
import glob

path = input("Enter the file path: ")

pattern = path + "\\*.pdf"
result = glob.glob(pattern)

for file_name in result:
   reader = PdfReader(file_name)
   page = reader.pages[0]
   text = page.extract_text()
   print(text)

提取出的文本大概是这样:

Payslip 
March 2024  
 randomtext 
123456Name of the employee

后面还有一堆无关内容。我的需求是每月从这些PDF里提取固定位置、4-6位长度的员工ID,用ID重命名文件。文件格式固定,但月份年份会变。之前试过字符串切片,因为月份变化、ID长度不固定搞砸了。

另外想问:有没有更好用的PDF读取库?怎么精准提取这个固定位置的可变长度ID?我自己写了一段代码,如下:

from pypdf import PdfReader
import os
import glob
import re

month = input("Which month's reports? Type the full name of the month: ")
"(In the following format: MM_YYYY) "
path = input("Enter the file path: ")

pattern = path + "\\*.pdf"
result = glob.glob(pattern)

try:
   for file_name in result:
      reader = PdfReader(file_name)
      page = reader.pages[0]
      text = page.extract_text()
      mylist = text.split()
      myIds = []
      myIds.append(mylist[4])
      mystring = ''.join(myIds)
      temp = re.findall(r'\d+', mystring)
      res = list(map(int, temp))
      old_name = file_name
      number = str(res)
      number = number.replace('[', '')
      number = number.replace(']', '')
      new_name = path + '\\' + number + "Company name Payslip" + "_" + \
      month + old_name[-4:]
      os.rename(old_name, new_name)
except FileExistsError:
   print("There is a duplicate payslip for EE ID " + number + " Delete one of their\
 payslips.")

解决思路

一、PDF读取库推荐

  • PyMuPDF(fitz):文本提取准确率比pypdf高不少,速度快,对复杂排版的PDF兼容性更好,很多场景下能避免文本乱序或丢失的问题。
  • PDFplumber:优势是能精准获取文本的坐标位置,如果你需要严格按固定区域提取内容,这个库很合适,还支持提取表格数据。

二、精准提取员工ID的方法

根据你的PDF格式,员工ID是固定位置的4-6位数字,给你两种可靠的提取方式:

方式1:正则匹配目标数字段

从提取的文本来看,ID紧跟在randomtext行之后,和员工姓名连在一起。直接用正则匹配这段里的4-6位数字就行:

import re

text = """Payslip 
March 2024  
 randomtext 
123456Name of the employee"""

# 匹配4-6位数字,且后面紧跟字母(员工姓名开头)
employee_id = re.search(r'\d{4,6}(?=[A-Za-z])', text).group()
print(employee_id)  # 输出123456

方式2:按行定位(更稳定)

因为PDF格式固定,ID所在的行位置是固定的(示例里是第4行),可以按行分割文本后提取:

# 按行分割并过滤空行
lines = [line.strip() for line in text.split('\n') if line.strip()]
# 取包含ID的行(示例里是第4行,索引为3)
id_line = lines[3]
# 提取行内的4-6位数字
employee_id = re.search(r'\d{4,6}', id_line).group()
print(employee_id)

三、优化后的重命名代码

结合上面的方法,优化你的代码,解决原代码里的索引依赖、正则处理冗余、路径兼容问题:

from pypdf import PdfReader
import os
import glob
import re

month = input("请输入月份(格式:MM_YYYY):")
path = input("请输入文件夹路径:")

# 用os.path.join处理路径,避免Windows/Linux分隔符问题
pattern = os.path.join(path, "*.pdf")
result = glob.glob(pattern)

for file_name in result:
    try:
        reader = PdfReader(file_name)
        page = reader.pages[0]
        text = page.extract_text()
        
        # 分割文本为行,过滤空行
        lines = [line.strip() for line in text.split('\n') if line.strip()]
        # 定位到ID所在行(根据你的PDF格式调整索引)
        if len(lines) < 4:
            print(f"文件 {file_name} 内容格式异常,跳过")
            continue
        id_line = lines[3]
        
        # 提取4-6位数字ID
        match = re.search(r'\d{4,6}', id_line)
        if not match:
            print(f"文件 {file_name} 未找到有效员工ID,跳过")
            continue
        employee_id = match.group()
        
        # 构造新文件名
        new_name = os.path.join(path, f"{employee_id}_Company name Payslip_{month}.pdf")
        
        # 检查文件是否已存在,避免重命名冲突
        if os.path.exists(new_name):
            print(f"员工ID {employee_id} 的工资单已存在,请删除重复文件后重试")
            continue
        
        os.rename(file_name, new_name)
        print(f"重命名成功:{os.path.basename(file_name)} -> {os.path.basename(new_name)}")
    except Exception as e:
        print(f"处理文件 {file_name} 时出错:{str(e)}")

优化点说明

  1. 用os.path.join处理路径,适配不同操作系统;
  2. 增加了格式校验,避免因PDF内容异常导致的索引错误;
  3. 提前检查文件是否存在,主动提示重复问题,不用靠异常捕获;
  4. 简化正则处理逻辑,直接提取目标数字;
  5. 用f-string构造文件名,代码更简洁易读;
  6. 增加详细的日志提示,方便排查问题。

内容的提问来源于stack exchange,提问作者Mr Cs

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 10:14:53