如何用Python(pandas)从PDF简历提取数据至CSV及OCR适配问题
PDF简历数据提取问题排查与解决方案
问题核心
你尝试用正则从PDF简历提取数据生成CSV,但得到空文件,同时需要将PDF转换为适合OCR的格式。
空文件/空数据的常见原因
- PDF类型不匹配:如果是扫描生成的图片型PDF,
pdfplumber无法提取文本,导致后续正则无匹配内容。 - 文本清理规则错误:原代码的
clean_text会删除所有中文字符,若简历是中文格式,标签(如“姓名”“电话”)和内容会被清空。 - 正则依赖固定英文标签:多数中文简历用“姓名”“邮箱”而非“Name”“Email”,导致正则完全匹配不到内容。
- 目录路径或文件问题:目标目录无PDF文件,或路径权限不足,导致循环未执行,
data列表为空。
分步解决方案
一、将PDF转换为适合OCR的格式(针对图片型PDF)
如果你的PDF是扫描件,先转成图片格式再做OCR:
from pdf2image import convert_from_path import os # PDF转图片 def pdf_to_images(pdf_path, output_img_dir): if not os.path.exists(output_img_dir): os.makedirs(output_img_dir) # Windows需提前安装poppler并配置路径,Mac/Linux可直接通过包管理器安装 images = convert_from_path(pdf_path, poppler_path=r"C:\poppler-24.02.0\Library\bin") for idx, img in enumerate(images): img_name = f"{os.path.splitext(os.path.basename(pdf_path))[0]}_page{idx+1}.jpg" img.save(os.path.join(output_img_dir, img_name), "JPEG") # 对图片做OCR提取文本 import pytesseract from PIL import Image def ocr_from_image(image_path): # 需安装Tesseract并配置环境变量,添加中文语言包 pytesseract.pytesseract.tesseract_cmd = r"C:\Program Files\Tesseract-OCR\tesseract.exe" text = pytesseract.image_to_string(image_path, lang="chi_sim+eng") return text
二、修复文本提取与正则匹配逻辑
- 修正文本清理函数:保留中文字符,避免关键内容丢失
def clean_text(text): # 保留中文、英文、数字、常用符号及空格 cleaned_text = re.sub(r"[^\u4e00-\u9fa5a-zA-Z0-9\s@.-]", "", text) cleaned_text = re.sub(r"\s+", " ", cleaned_text) return cleaned_text.strip()
- 优化正则规则:支持中英文标签,增加灵活匹配
def extract_data_from_resume(pdf_path): text = "" # 先尝试直接提取PDF文本 with pdfplumber.open(pdf_path) as pdf: for page in pdf.pages: page_text = page.extract_text() if page_text: text += page_text # 如果无文本提取结果,转图片做OCR if not text: temp_img_dir = "temp_imgs" pdf_to_images(pdf_path, temp_img_dir) for img_file in os.listdir(temp_img_dir): if img_file.endswith(".jpg"): text += ocr_from_image(os.path.join(temp_img_dir, img_file)) # 清理临时文件 for img_file in os.listdir(temp_img_dir): os.remove(os.path.join(temp_img_dir, img_file)) os.rmdir(temp_img_dir) cleaned_text = clean_text(text) # 适配中英文标签的正则 name_pattern = r"(?:姓名|Name):?\s*(.*?)(?:\s+|电话|Phone|邮箱|Email|$)" name_match = re.search(name_pattern, cleaned_text, re.IGNORECASE) name = name_match.group(1).strip() if name_match else "" email_pattern = r"(?:邮箱|Email):?\s*([^\s@]+@[^\s@]+\.[^\s@]+)" email_match = re.search(email_pattern, cleaned_text, re.IGNORECASE) email = email_match.group(1).strip() if email_match else "" phone_pattern = r"(?:电话|Phone):?\s*([\d\-\+\s]{8,20})" phone_match = re.search(phone_pattern, cleaned_text, re.IGNORECASE) phone = phone_match.group(1).strip() if phone_match else "" return name, email, phone
- 增加路径校验:确保目录有效且存在PDF文件
def convert_pdf_to_csv(directory, output_csv_path): data = [] # 校验目录合法性 if not os.path.isdir(directory): print(f"错误:目录 {directory} 不存在") return # 遍历处理PDF文件 for filename in os.listdir(directory): if filename.lower().endswith(".pdf"): pdf_path = os.path.join(directory, filename) print(f"正在处理:{filename}") name, email, phone = extract_data_from_resume(pdf_path) data.append({"Name": name, "Email": email, "Phone": phone}) if not data: print("警告:未找到可处理的PDF文件,或所有文件未提取到有效数据") df = pd.DataFrame(data) # 用utf_8_sig避免中文乱码 df.to_csv(output_csv_path, index=False, encoding="utf_8_sig") print(f"CSV文件已生成:{output_csv_path}")
注意事项
- 需提前安装依赖包:
pip install pdfplumber pandas pdf2image pytesseract pillow poppler和Tesseract-OCR需手动安装并配置系统环境变量,或在代码中指定路径- 正则规则可根据你的简历格式进一步调整,比如适配不同的电话格式、姓名排列方式
内容的提问来源于stack exchange,提问作者user16001720
相关产品推荐
相关产品推荐

