You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python从PDF提取数据转CSV:日期与Patient ID提取失败求助

问题:PDF字段提取与CSV生成异常解决

问题概述

使用Python结合pdfplumber、pandas、re库从指定PDF提取Design File Name、Order Creation Date、Patient ID等字段并生成CSV时,遇到以下问题:

  • 无法提取Order Creation Date和Patient ID字段值
  • 生成的CSV数据分散在多行,大部分行值为NaN

环境依赖

  • 安装命令:pip install pdfplumber

当前代码

import pandas as pd
import pdfplumber
import re
from datetime import datetime  # Import the datetime module

# Defined PDF file path
pdf_path = "/content/drive/MyDrive/Final Project PDFs/Zoho Creator - Print              Queue Report.pdf"

# Open the PDF file
with pdfplumber.open(pdf_path) as pdf:
    # Extract text from each page
    text = ""
    for page in pdf.pages:
        text += page.extract_text()

    # Define a pattern to extract the Order Creation Date
order_creation_date_pattern = r"Order Creation Date (from New Cast Scan page).* (.+(\d+-[A-Z].*-\d+))"

# Find the Order Creation Date in the text
match = re.search(order_creation_date_pattern, text)
order_creation_date = match.group(1) if match else None

# Defined dictionary to map column names to their corresponding patterns
column_patterns = {
"Design File Name": r"Design File Name (.+)",
"Order Creation Date": r"Order Creation Date (from New Cast Scan page).* (.+(\d+-[A-Z].*-\d+))",
"Clinic Name": r"Clinic Name (.+)",
"Patient ID": r"Patient ID (.+)",
"Name on Shipping Label": r"Name on Shipping Label (.+)",
"Include Bungee Conversion Kit": r"Include Bungee Conversion Kit? (.+)",
}

# Add the "Closure" pattern
column_patterns["Closure"] = r"Closure Easily Removed (Splint) (.+)"

# Extract data based on the defined patterns
data = {}
for column, pattern in column_patterns.items():
data[column] = [match.group(1) if (match := re.search(pattern, text)) else None for text in text.split('\n')]

# Add the Order Creation Date to the data dictionary
data["Order Creation Date"] = [order_creation_date] * len(data[next(iter(data))])

# Creates a DataFrame from the extracted data
result_df = pd.DataFrame(data)

# Exports the DataFrame to a CSV file
result_df.to_csv("output.csv", index=False)

import pandas as pd

# Reads the CSV file
df = pd.read_csv("output.csv")

# Displays the DataFrame
print(df)

当前输出

Design File Name  Order Creation Date         Clinic Name  
0                          NaN                  NaN                 NaN
1   AA2_HT_XX_RD_A_B_MON_410_8                  NaN                 NaN
2                          NaN                  NaN                 NaN
...(省略其余行)

Patient ID Name on Shipping Label Include Bungee Conversion Kit  Closure  
0          NaN                    NaN                           NaN      NaN
1          NaN                    NaN                           NaN      NaN
...(省略其余行)

期望输出

生成仅包含单条数据的CSV表格,所有字段对应值在同一行,无多余NaN行。

PDF样本内容

Print Queue Report
Design File Name AA2Redo_HT_XX_RD_A_B_MON_410_8
Design Creation Date 08-Apr-2024
Order Creation Date (from New Cast Scan page) 08-Apr-2024
Shipping Comments
Clinic Name Monarch Wound Care
Custom Clinic Logo No
Custom Clinic Logo Text
Patient ID AA2Redo
...(省略其余内容)


解决方案

1. 修复正则表达式匹配问题

原正则存在两个核心问题:

  • Order Creation Date的正则未转义括号,导致无法匹配带括号的文本;冗余分组干扰结果提取
  • 部分字段正则未考虑无值场景,容错性不足

修正后的正则字典:

column_patterns = {
    "Design File Name": r"Design File Name (.+)",
    "Order Creation Date": r"Order Creation Date \(from New Cast Scan page\) (.+)",
    "Clinic Name": r"Clinic Name (.+)",
    "Patient ID": r"Patient ID (.+)",
    "Name on Shipping Label": r"Name on Shipping Label (.+)?",  # 允许字段无值
    "Include Bungee Conversion Kit": r"Include Bungee Conversion Kit\? (.+)",
    "Closure": r"Closure Easily Removed \(Splint\) (.+)?"
}

2. 修正数据提取逻辑

原代码错误地遍历每行文本重复匹配,生成与行数一致的列表,导致多行NaN。正确逻辑应为:在整个PDF文本中一次性提取每个字段的值,构造单条数据字典。

3. 完整修正代码

import pandas as pd
import pdfplumber
import re

# PDF文件路径
pdf_path = "/content/drive/MyDrive/Final Project PDFs/Zoho Creator - Print              Queue Report.pdf"

# 提取PDF文本
with pdfplumber.open(pdf_path) as pdf:
    text = ""
    for page in pdf.pages:
        text += page.extract_text()

# 定义字段与正则映射
column_patterns = {
    "Design File Name": r"Design File Name (.+)",
    "Order Creation Date": r"Order Creation Date \(from New Cast Scan page\) (.+)",
    "Clinic Name": r"Clinic Name (.+)",
    "Patient ID": r"Patient ID (.+)",
    "Name on Shipping Label": r"Name on Shipping Label (.+)?",
    "Include Bungee Conversion Kit": r"Include Bungee Conversion Kit\? (.+)",
    "Closure": r"Closure Easily Removed \(Splint\) (.+)?"
}

# 提取单条数据
data = {}
for column, pattern in column_patterns.items():
    match = re.search(pattern, text)
    data[column] = match.group(1).strip() if match else None

# 生成单行DataFrame并导出CSV
result_df = pd.DataFrame([data])
result_df.to_csv("output.csv", index=False)

# 验证输出
df = pd.read_csv("output.csv")
print(df)

4. 效果说明

  • 修正后的代码会在整个PDF文本中精准匹配每个字段,提取单个有效值
  • 生成的DataFrame仅包含一行有效数据,导出的CSV无多余NaN行
  • 所有目标字段(包括日期、Patient ID)均可正确提取

内容的提问来源于stack exchange,提问作者kevin marquez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 09:02:09