求助:构建多规则正则表达式优化医疗转录文本处理
Hey there, brute-force processing for medical transcription text sounds tedious—regex is exactly the tool to streamline this. Let’s break down each of your formatting rules into targeted regex patterns, explain how they work, and include sample inputs/outputs to validate the solution.
Key Rules & Regex Implementations
All patterns assume multi-line mode is enabled (e.g., re.MULTILINE in Python, /gm flag in JavaScript) to treat each line as a separate match context.
Rule 1: Isolated camelCase/lowercase lines → UPPERCASE with colon
Target lines that are camelCase or lowercase, have no trailing colon, and are surrounded by empty lines.
Regex Pattern:
(?<=^\n)(^[a-z][a-zA-Z]*$)(?=\n^$)
Replacement:
- Python:
\U\1:(uses uppercase conversion flag) - JavaScript:
(match) => match.toUpperCase() + ":"
Explanation: (?<=^\n): Ensures the line is preceded by an empty line^[a-z][a-zA-Z]*$: Matches camelCase or lowercase word chains with no colon(?=\n^$): Ensures the line is followed by an empty line
Rule 2: UPPERCASE line without colon followed by regular text → add colon
Target uppercase lines missing a colon, where the next line is regular transcribed text.
Regex Pattern:
(^[A-Z]+$)(\n)(^[a-z].*)
Replacement: \1:\2\3
Explanation:
- Captures the uppercase line, the newline, and the following regular text
- Adds a colon to the end of the uppercase line while preserving the rest of the content
Rule 3: UPPERCASE line without colon followed by colon-containing line → replace with space
Target uppercase lines that are redundant because the next line already has a labeled heading.
Regex Pattern:
^[A-Z]+$\n(?=^[a-zA-Z]+:)
Replacement: (single space)
Explanation:
- Matches the uppercase line and its trailing newline
- Uses a lookahead to confirm the next line has a labeled heading (ends with colon)
- Replaces the redundant line with a space to keep the following content properly formatted
Additional Replacements
These handle the extra cleanup tasks you mentioned:
- Replace separator lines:
Pattern:\n\s*\*\s*\*\s*\*\s*\n
Replacement: - Remove sample report label:
Pattern:\(Medical Transcription Sample Report\)
Replacement:
Sample Input/Output Examples
Sample 1 (Rule 1 Application)
Input:
historyOfPresentIllness The patient reports experiencing chest pain for 3 days.
Output:
HISTORYOFPRESENTILLNESS: The patient reports experiencing chest pain for 3 days.
Sample 2 (Rule 2 Application)
Input:
MEDICALHISTORY Patient has a history of hypertension.
Output:
MEDICALHISTORY: Patient has a history of hypertension.
Sample 3 (Rule 3 + Cleanup Application)
Input:
CURRENTMEDICATIONS Aspirin: 81mg daily * * * (Medical Transcription Sample Report)
Output:
Aspirin: 81mg daily
Note: For best results, apply the cleanup replacements first, then Rule 3, Rule 2, and finally Rule 1. Test with your specific dataset to adjust for any edge cases in transcription formatting.
内容的提问来源于stack exchange,提问作者user2535338

