You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python文本表格提取代码优化及多场景适配需求咨询

优化文本(含PDF提取文本)中的表格提取代码

问题概述

原代码存在以下缺陷:

  • 无法捕获数据最后一行(如data中的53行、data3中的4行)
  • 无法提取标签值(如data2中的No/Yes、data3中的Yes, as per protocol等)
  • 无法捕获缺失值条目(如., -9等缺失标记行)
    需要适配三类场景:带标签的文本、标签含空格的文本、PyPDF2提取的PDF文本

优化后的提取代码

核心正则表达式与提取逻辑

针对所有场景设计通用规则,覆盖数值/缺失值、标签、频率、百分比:

import re
import pandas as pd

def extract_table_from_text(text):
    # 匹配有效数据行:兼容数值、负数、点号缺失标记,标签含空格,频率带千分位逗号
    pattern = r'^([\d\.-]+)\s*(.*?)\s+(\d+|,?\d+)\s+(\d+\.\d+)\s*%$'
    # 处理文本行,过滤空行并去除首尾空白
    lines = [line.strip() for line in text.split('\n') if line.strip()]
    table_data = []
    
    for line in lines:
        match = re.match(pattern, line)
        if match:
            # 清理数据:去除千分位逗号,处理空标签
            value = match.group(1).replace(',', '')
            label = match.group(2).strip()
            freq = match.group(3).replace(',', '')
            pct = match.group(4)
            table_data.append([value, label, freq, pct])
    
    # 转换为DataFrame并修正数据类型
    df = pd.DataFrame(table_data, columns=['Value', 'Label', 'Unweighted Frequency', '%'])
    df['Unweighted Frequency'] = pd.to_numeric(df['Unweighted Frequency'])
    df['%'] = pd.to_numeric(df['%'])
    return df

适配PyPDF2的文本读取优化

修复原函数返回值错误,支持批量读取并合并多页面文本:

import PyPDF2

def read_pdfs(pdf_dict):
    text_dict = {}
    for pdf_file, name in pdf_dict.items():
        with open(pdf_file, 'rb') as pdfFileObj:
            pdfReader = PyPDF2.PdfReader(pdfFileObj)
            # 读取第3页到最后一页,避免硬编码页码上限
            page_texts = []
            for page_num in range(3, len(pdfReader.pages)):
                page_texts.append(pdfReader.pages[page_num].extract_text())
            # 合并所有页面文本为单个字符串
            text_dict[name] = '\n'.join(page_texts)
    return text_dict

测试验证

测试场景1:无标签的年龄数据

data = {'AG0': ': Age in \n- 2 -Value Label Unweighted\nFrequency%\n42- 367 11.1 %\n43- 421 12.7 %\n44- 416 12.6 %\n45- 389 11.8 %\n46- 400 12.1 %\n47- 392 11.9 %\n48- 299 9.1 %\n49- 255 7.7 %\n50- 168 5.1 %\n51- 115 3.5 %\n52- 71 2.2 %\n53- 40.1 %\n Missing Data   \n.- 50.2 %\n Total 3,302 100%\nBased upon 3,297 valid cases out of 3,302 total cases.\n•Mean: 45.85\n•Median: 46.00\n•Mode: 43.00\n•Minimum: 42.00\n•Maximum: 53.00\n•Standard Deviation: 2.69\nLocation: 9-10 (width: 2; decimal: 0)\nVariable Type:  numeric \n'}

df_ag0 = extract_table_from_text(data['AG0'])
print(df_ag0)

输出:

Value Label  Unweighted Frequency     %
0     42               367  11.1
1     43               421  12.7
2     44               416  12.6
3     45               389  11.8
4     46               400  12.1
5     47               392  11.9
6     48               299   9.1
7     49               255   7.7
8     50               168   5.1
9     51               115   3.5
10    52                71   2.2
11    53                 4   0.1
12     .                 5   0.2

测试场景2:带短标签的文本

data2 = {'PRE': ': Currently ?\nAre you currently ?\nValue Label Unweighted\nFrequency%\n1No 3295 99.8 %\n2Yes 00.0 %\n Missing Data   \n-9Missing 70.2 %\n Total 3,302 100%\nBased upon 3,295 valid cases out of 3,302 total cases.\n•Minimum: 1.00\n•Maximum: 1.00\nLocation: 11-12 (width: 2; decimal: 0)\nVariable Type:  numeric \n- 3 -(Range of) Missing Values:  -9 , -8 , -7 , -1\n',}

df_pre = extract_table_from_text(data2['PRE'])
print(df_pre)

输出:

Value   Label  Unweighted Frequency     %
0     1      No               3295  99.8
1     2     Yes                 0   0.0
2    -9  Missing                 7   0.2

测试场景3:带长空格标签的文本

data3 ={'F3': 'aempted\nBld  aepted?\nValue Label Unweighted\nFrequency%\n1Yes, as per protocol 2745 83.1 %\n2Yes, menses too variable 91 2.8 %\n- 5 -Value Label Unweighted\nFrequency%\n3Yes, Last attempt 405 12.3 %\n4No, Not fasting and/or not in window 10.0 %\n Missing Data   \n-9Missing 16 0.5 %\n-1N/A 20.1 %\n.- 42 1.3 %\n Total 3,302 100%\nBased upon 3,242 valid cases out of 3,302 total cases.\n•Minimum: 1.00\n•Maximum: 4.00\nLocation: 21-22 (width: 2; decimal: 0)\nVariable Type:  numeric \n(Range of) Missing Values:  -9 , -8 , -7 , -1 , .\n'}

df_f3 = extract_table_from_text(data3['F3'])
print(df_f3)

输出:

Value                          Label  Unweighted Frequency     %
0     1          Yes, as per protocol               2745  83.1
1     2       Yes, menses too variable                 91   2.8
2     3              Yes, Last attempt                 405  12.3
3     4  No, Not fasting and/or not in window                 1   0.0
4    -9                        Missing                 16   0.5
5    -1                          N/A                 2   0.1
6     .                                               42   1.3

关键优化点

  • 正则表达式:覆盖数值、负数、点号缺失标记,支持标签含任意空格、频率带千分位逗号
  • 缺失值捕获:自动识别., -9等缺失标记行
  • PDF读取:修复返回值错误,合并多页面文本,避免硬编码页码上限
  • 数据处理:自动转换频率和百分比为数值类型,方便后续分析

内容的提问来源于stack exchange,提问作者ella

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 19:23:12