Python批量提取病理报告TXT指定字段的优化方案问询
Hey there! Let's figure out how to efficiently extract those structured fields from your 500+ pathology report TXT files and get them into the array format you need.
Approach Overview
Since your reports have consistent labeled fields (Case, Sex, Age, etc.), regular expressions are perfect here—they’re flexible enough to handle minor formatting variations (like extra spaces or capitalization differences) and fast enough to process hundreds of files quickly. We’ll batch-process all files in your target folder, extract each field in your desired order, and compile everything into an array of arrays.
Step-by-Step Implementation (Python)
Python is ideal for this task thanks to its built-in file handling and regex support. Here’s a complete, reusable script:
import os import re # Define regex patterns for each field—tweak these if your TXT format has slight variations FIELD_PATTERNS = { 'Case': re.compile(r'Case:\s*(.*)', re.IGNORECASE), 'Sex': re.compile(r'Sex:\s*(.*)', re.IGNORECASE), 'Age': re.compile(r'Age:\s*(.*)', re.IGNORECASE), 'COLLECTED': re.compile(r'COLLECTED:\s*(.*)', re.IGNORECASE), 'REPORTED': re.compile(r'REPORTED:\s*(.*)', re.IGNORECASE), 'DIAGNOSIS': re.compile(r'DIAGNOSIS:\s*(.*)', re.DOTALL) # DOTALL lets . match newlines for multi-line diagnoses } def extract_single_report(file_path): """Extract fields from a single pathology report TXT""" with open(file_path, 'r', encoding='utf-8') as f: content = f.read() # Follow your exact desired output order output_order = ['Case', 'Sex', 'Age', 'COLLECTED', 'REPORTED', 'DIAGNOSIS'] extracted_values = [] for field in output_order: match = FIELD_PATTERNS[field].search(content) if match: # Clean up the value—strip extra whitespace, and flatten multi-line diagnoses raw_value = match.group(1) cleaned_value = raw_value.strip() if field != 'DIAGNOSIS' else raw_value.strip().replace('\n', ' ') extracted_values.append(cleaned_value) else: # Handle missing fields (replace with empty string or None if preferred) extracted_values.append('') return extracted_values def process_all_reports(folder_path): """Batch process all TXT files in the target folder""" all_reports_data = [] for filename in os.listdir(folder_path): if filename.lower().endswith('.txt'): full_path = os.path.join(folder_path, filename) report_data = extract_single_report(full_path) all_reports_data.append(report_data) print(f"Processed: {filename}") # Optional progress update return all_reports_data # Run the script if __name__ == "__main__": # Replace with your actual folder path containing the TXT files reports_directory = "./pathology_reports" final_result = process_all_reports(reports_directory) # Example: Print the first 2 entries to verify print("\nSample Output:") for entry in final_result[:2]: print(entry)
Key Details to Note
- Regex Flexibility: The
re.IGNORECASEflag ensures the script works even if fields are capitalized differently (e.g.,case:instead ofCase:). - Multi-Line Diagnoses: The
re.DOTALLflag for DIAGNOSIS lets it capture the entire multi-line diagnostic text, which we then flatten into a single string with spaces instead of newlines (matching your example output). - Missing Fields: If a report is missing a field, the script adds an empty string (you can swap this for
Noneif you prefer to explicitly mark missing values). - Efficiency: This script will handle 500 files quickly—each file is read once, and regex searches are optimized for speed.
Customization Tips
If your TXT files have a different format (e.g., fields are separated by colons without labels, or values are on the next line), just adjust the regex patterns to match your actual file structure. For example, if Age is always on a line starting with Age:, the existing pattern will work perfectly.
内容的提问来源于stack exchange,提问作者Peter S

