Python医疗账单文本提取:不规则跨行日期格式处理问题
解决医疗账单文本跨行日期数据提取与Excel导出问题
原代码仅能处理单行格式的医疗账单数据,无法应对部分账单的跨行结构——这类账单的起始日期、服务类型、金额信息在第一行,结束日期与程序编号在第二行,导致数据提取不全。以下是修正后的解决方案:
修正后的代码
import pandas as pd import re input_file_path = r'C:\Users\test\Downloads\PracticalAssessmentFiles\Input.txt' output_file_path = r'C:\Users\test\Downloads\PracticalAssessmentFiles\output.xlsx' # 按行读取并预处理文本(去除空行与首尾空格) with open(input_file_path, 'r') as file: lines = [line.strip() for line in file if line.strip()] processed_lines = [] i = 0 while i < len(lines): current_line = lines[i] # 识别需要合并的跨行条目:当前行包含起始日期、服务类型和所有金额字段,但缺少程序编号 if re.match(r'\d{2}/\d{2}/\d{2,4}\s+[\w\s]+\s+[\d.]+\s+[\d.]+\s+[\d.]+\s+[\d.]+\s+[\d.]+\s+[\d.]+$', current_line): if i + 1 < len(lines): # 合并当前行与下一行的结束日期、程序编号信息 merged_line = f"{current_line} {lines[i+1]}" processed_lines.append(merged_line) i += 2 continue processed_lines.append(current_line) i += 1 # 合并处理后的行,执行正则匹配 input_string = '\n'.join(processed_lines) pattern = r'(\d{2}/\d{2}/\d{2,4}(?:\s*-\s*\d{2}/\d{2}/\d{2,4})?)\s+([\w\s]+)\s+([\dA-Z]+)\s+([\d.]+)\s+([\d.]+)\s+([\d.]+)\s+([\d.]+)\s+([\d.]+)\s+([\d.]+)' matches = re.findall(pattern, input_string) # 转换为结构化数据 data = [] for match in matches: service_date, service_type, procedure_number, billed, allowed, paid, non_covered, deductible, copayment = match data.append({ "Service Date": service_date, "Type of Service": service_type, "Procedure Number": procedure_number, "Amount Billed": billed, "Amount Allowed": allowed, "Amount We Paid": paid, "Non Covered": non_covered, "Deductible": deductible, "Copayment": copayment }) # 导出到Excel文件 df = pd.DataFrame(data) df.to_excel(output_file_path, index=False) print(f"数据已处理并保存至 {output_file_path}")
关键逻辑说明
- 跨行合并处理:遍历文本行,通过正则识别出不完整的账单行,将其与下一行合并,确保每个账单的所有信息都在同一行中。
- 正则匹配兼容:保留原正则的匹配规则,同时支持单日期格式和合并后的双日期格式(起始日期-结束日期)。
- 结构化数据导出:将提取到的字段映射为指定的Excel列名,最终导出为无索引的Excel文件。
内容的提问来源于stack exchange,提问作者prudhvi
相关产品推荐
相关产品推荐

