Python从PDF提取数据转CSV:日期与Patient ID提取失败求助
问题:PDF字段提取与CSV生成异常解决
问题概述
使用Python结合pdfplumber、pandas、re库从指定PDF提取Design File Name、Order Creation Date、Patient ID等字段并生成CSV时,遇到以下问题:
- 无法提取
Order Creation Date和Patient ID字段值 - 生成的CSV数据分散在多行,大部分行值为NaN
环境依赖
- 安装命令:
pip install pdfplumber
当前代码
import pandas as pd import pdfplumber import re from datetime import datetime # Import the datetime module # Defined PDF file path pdf_path = "/content/drive/MyDrive/Final Project PDFs/Zoho Creator - Print Queue Report.pdf" # Open the PDF file with pdfplumber.open(pdf_path) as pdf: # Extract text from each page text = "" for page in pdf.pages: text += page.extract_text() # Define a pattern to extract the Order Creation Date order_creation_date_pattern = r"Order Creation Date (from New Cast Scan page).* (.+(\d+-[A-Z].*-\d+))" # Find the Order Creation Date in the text match = re.search(order_creation_date_pattern, text) order_creation_date = match.group(1) if match else None # Defined dictionary to map column names to their corresponding patterns column_patterns = { "Design File Name": r"Design File Name (.+)", "Order Creation Date": r"Order Creation Date (from New Cast Scan page).* (.+(\d+-[A-Z].*-\d+))", "Clinic Name": r"Clinic Name (.+)", "Patient ID": r"Patient ID (.+)", "Name on Shipping Label": r"Name on Shipping Label (.+)", "Include Bungee Conversion Kit": r"Include Bungee Conversion Kit? (.+)", } # Add the "Closure" pattern column_patterns["Closure"] = r"Closure Easily Removed (Splint) (.+)" # Extract data based on the defined patterns data = {} for column, pattern in column_patterns.items(): data[column] = [match.group(1) if (match := re.search(pattern, text)) else None for text in text.split('\n')] # Add the Order Creation Date to the data dictionary data["Order Creation Date"] = [order_creation_date] * len(data[next(iter(data))]) # Creates a DataFrame from the extracted data result_df = pd.DataFrame(data) # Exports the DataFrame to a CSV file result_df.to_csv("output.csv", index=False) import pandas as pd # Reads the CSV file df = pd.read_csv("output.csv") # Displays the DataFrame print(df)
当前输出
Design File Name Order Creation Date Clinic Name 0 NaN NaN NaN 1 AA2_HT_XX_RD_A_B_MON_410_8 NaN NaN 2 NaN NaN NaN ...(省略其余行) Patient ID Name on Shipping Label Include Bungee Conversion Kit Closure 0 NaN NaN NaN NaN 1 NaN NaN NaN NaN ...(省略其余行)
期望输出
生成仅包含单条数据的CSV表格,所有字段对应值在同一行,无多余NaN行。
PDF样本内容
Print Queue Report
Design File Name AA2Redo_HT_XX_RD_A_B_MON_410_8
Design Creation Date 08-Apr-2024
Order Creation Date (from New Cast Scan page) 08-Apr-2024
Shipping Comments
Clinic Name Monarch Wound Care
Custom Clinic Logo No
Custom Clinic Logo Text
Patient ID AA2Redo
...(省略其余内容)
解决方案
1. 修复正则表达式匹配问题
原正则存在两个核心问题:
Order Creation Date的正则未转义括号,导致无法匹配带括号的文本;冗余分组干扰结果提取- 部分字段正则未考虑无值场景,容错性不足
修正后的正则字典:
column_patterns = { "Design File Name": r"Design File Name (.+)", "Order Creation Date": r"Order Creation Date \(from New Cast Scan page\) (.+)", "Clinic Name": r"Clinic Name (.+)", "Patient ID": r"Patient ID (.+)", "Name on Shipping Label": r"Name on Shipping Label (.+)?", # 允许字段无值 "Include Bungee Conversion Kit": r"Include Bungee Conversion Kit\? (.+)", "Closure": r"Closure Easily Removed \(Splint\) (.+)?" }
2. 修正数据提取逻辑
原代码错误地遍历每行文本重复匹配,生成与行数一致的列表,导致多行NaN。正确逻辑应为:在整个PDF文本中一次性提取每个字段的值,构造单条数据字典。
3. 完整修正代码
import pandas as pd import pdfplumber import re # PDF文件路径 pdf_path = "/content/drive/MyDrive/Final Project PDFs/Zoho Creator - Print Queue Report.pdf" # 提取PDF文本 with pdfplumber.open(pdf_path) as pdf: text = "" for page in pdf.pages: text += page.extract_text() # 定义字段与正则映射 column_patterns = { "Design File Name": r"Design File Name (.+)", "Order Creation Date": r"Order Creation Date \(from New Cast Scan page\) (.+)", "Clinic Name": r"Clinic Name (.+)", "Patient ID": r"Patient ID (.+)", "Name on Shipping Label": r"Name on Shipping Label (.+)?", "Include Bungee Conversion Kit": r"Include Bungee Conversion Kit\? (.+)", "Closure": r"Closure Easily Removed \(Splint\) (.+)?" } # 提取单条数据 data = {} for column, pattern in column_patterns.items(): match = re.search(pattern, text) data[column] = match.group(1).strip() if match else None # 生成单行DataFrame并导出CSV result_df = pd.DataFrame([data]) result_df.to_csv("output.csv", index=False) # 验证输出 df = pd.read_csv("output.csv") print(df)
4. 效果说明
- 修正后的代码会在整个PDF文本中精准匹配每个字段,提取单个有效值
- 生成的DataFrame仅包含一行有效数据,导出的CSV无多余NaN行
- 所有目标字段(包括日期、Patient ID)均可正确提取
内容的提问来源于stack exchange,提问作者kevin marquez
相关产品推荐
相关产品推荐

