You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:构建多规则正则表达式优化医疗转录文本处理

Solution for Medical Transcription Text Processing with Regex

Hey there, brute-force processing for medical transcription text sounds tedious—regex is exactly the tool to streamline this. Let’s break down each of your formatting rules into targeted regex patterns, explain how they work, and include sample inputs/outputs to validate the solution.


Key Rules & Regex Implementations

All patterns assume multi-line mode is enabled (e.g., re.MULTILINE in Python, /gm flag in JavaScript) to treat each line as a separate match context.

Rule 1: Isolated camelCase/lowercase lines → UPPERCASE with colon

Target lines that are camelCase or lowercase, have no trailing colon, and are surrounded by empty lines.

Regex Pattern:

(?<=^\n)(^[a-z][a-zA-Z]*$)(?=\n^$)

Replacement:

  • Python: \U\1: (uses uppercase conversion flag)
  • JavaScript: (match) => match.toUpperCase() + ":"
    Explanation:
  • (?<=^\n): Ensures the line is preceded by an empty line
  • ^[a-z][a-zA-Z]*$: Matches camelCase or lowercase word chains with no colon
  • (?=\n^$): Ensures the line is followed by an empty line

Rule 2: UPPERCASE line without colon followed by regular text → add colon

Target uppercase lines missing a colon, where the next line is regular transcribed text.

Regex Pattern:

(^[A-Z]+$)(\n)(^[a-z].*)

Replacement: \1:\2\3
Explanation:

  • Captures the uppercase line, the newline, and the following regular text
  • Adds a colon to the end of the uppercase line while preserving the rest of the content

Rule 3: UPPERCASE line without colon followed by colon-containing line → replace with space

Target uppercase lines that are redundant because the next line already has a labeled heading.

Regex Pattern:

^[A-Z]+$\n(?=^[a-zA-Z]+:)

Replacement: (single space)
Explanation:

  • Matches the uppercase line and its trailing newline
  • Uses a lookahead to confirm the next line has a labeled heading (ends with colon)
  • Replaces the redundant line with a space to keep the following content properly formatted

Additional Replacements

These handle the extra cleanup tasks you mentioned:

  1. Replace separator lines:
    Pattern: \n\s*\*\s*\*\s*\*\s*\n
    Replacement:
  2. Remove sample report label:
    Pattern: \(Medical Transcription Sample Report\)
    Replacement:

Sample Input/Output Examples

Sample 1 (Rule 1 Application)

Input:

historyOfPresentIllness

The patient reports experiencing chest pain for 3 days.

Output:

HISTORYOFPRESENTILLNESS:
The patient reports experiencing chest pain for 3 days.

Sample 2 (Rule 2 Application)

Input:

MEDICALHISTORY
Patient has a history of hypertension.

Output:

MEDICALHISTORY:
Patient has a history of hypertension.

Sample 3 (Rule 3 + Cleanup Application)

Input:

CURRENTMEDICATIONS
Aspirin: 81mg daily

* * *

(Medical Transcription Sample Report)

Output:

Aspirin: 81mg daily

 

Note: For best results, apply the cleanup replacements first, then Rule 3, Rule 2, and finally Rule 1. Test with your specific dataset to adjust for any edge cases in transcription formatting.

内容的提问来源于stack exchange,提问作者user2535338

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:45:01