You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则表达式匹配大写长单词及前后4个上下文单词的问题排查

Fixing Regex for Extracting Clinical Trial Identifiers with Context

Let's break down why your original regex is failing and how to fix it to reliably capture those 4+ letter all-caps trial IDs (like ASPIRE) along with their surrounding 4-word context.

What's Wrong with the Original Regex?

Your regex ((?:\w*\s*){,4})\s*([A-Z]{4,})\s*((?:\s*\w*){,4}) has a couple key issues:

  • No word boundaries: It can match partial words (e.g., if you have abcASPIRE, it would capture abc as part of the pre-context instead of recognizing ASPIRE as an independent identifier).
  • Poor punctuation handling: It doesn't account for commas, periods, or parentheses around the ID, leading to truncated or messy context captures.
  • Loose word matching: \w* can match empty strings or partial word fragments, making the 4-word count inconsistent.

Improved Regex Solution

This version addresses those gaps by focusing on full words, handling punctuation, and ensuring we only match standalone all-caps IDs:

((?:\b[\w'-]+\b[\s.,;:!?]*){0,4})\b([A-Z]{4,})\b((?:[\s.,;:!?]*\b[\w'-]+\b){0,4})

Breakdown of the Components:

  • Pre-Context Capture Group: ((?:\b[\w'-]+\b[\s.,;:!?]*){0,4})

    • \b[\w'-]+\b: Matches full words (supports apostrophes like patient's and hyphens like phase-II to cover common clinical terminology).
    • [\s.,;:!?]*: Accounts for spaces and punctuation that often follow words in documents.
    • {0,4}: Limits the pre-context to a maximum of 4 words (works even if the ID is at the start of the text, returning an empty pre-context).
  • Trial ID Capture Group: \b([A-Z]{4,})\b

    • \b: Ensures we only match the ID as a standalone word (avoids partial matches like ABCDE123).
    • [A-Z]{4,}: Targets exactly what you need—all-caps sequences of 4+ letters.
  • Post-Context Capture Group: ((?:[\s.,;:!?]*\b[\w'-]+\b){0,4})

    • Mirrors the pre-context but starts with punctuation/spaces, so it correctly captures words after commas or parentheses following the ID.
    • Again, limits to 4 words, even if the ID is at the end of the text.

Test Case Examples

Let's verify this works with common scenarios:

  1. ID with trailing punctuation:

    The primary endpoint of ASPIRE, a phase 3 oncology trial, met its targets

    • Pre-context: The primary endpoint of
    • ID: ASPIRE
    • Post-context: , a phase 3
  2. ID at the start of text:

    ENDEAVOR was designed to compare two first-line treatments

    • Pre-context: `` (empty)
    • ID: ENDEAVOR
    • Post-context: was designed to compare
  3. ID with surrounding parentheses:

    Data from (EXAMINE) supports the use of this therapy

    • Pre-context: Data from (
    • ID: EXAMINE
    • Post-context: ) supports the use

Fine-Tuning Options

  • If your trial IDs never include hyphens/apostrophes, you can simplify [\w'-]+ to \w+.
  • To exclude IDs that might accidentally include numbers (though rare for clinical trial acronyms), the \b boundary already prevents this since numbers count as word characters—ABCDE123 won't be matched.

内容的提问来源于stack exchange,提问作者Chris Daly

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:01:55