You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PowerShell正则提取名词短语失败,Pester测试返回Null求助

问题描述

我正尝试用PowerShell正则表达式,从带标签和属性的特定格式字符串中提取两类名词短语:

  • 由B-NP-singular与E-NP-singular配对组成的名词短语
  • 由B-NP-singular E-NP-singular B-PP B-NP-singular E-NP-singular组合成的名词短语

输入示例与预期结果

示例1

输入:

<S> A[a/DT,B-NP-singular] Whole[whole/JJ,whole/NN,whole/RP,I-NP-singular] New[New/JJ,New/NNP,new/JJ,I-NP-singular] Mind[mind/NN:UN,E-NP-singular] Why[why/CC,why/NN,why/UH,why/WRB,B-ADVP] Right-Brainers[Right-Brainers/null,B-NP-singular|E-NP-singular] Will[will/MD,B-VP] Rule[rule/VB,I-VP] the[the/DT,B-NP-singular] Future[future/JJ,future/NN:UN,</S>,E-NP-singular]<P/> 

预期提取:Whole New Mind、Right-Brainers、the Future

示例2

输入:

<S> Bit[bit/NN,B-NP-singular|E-NP-singular] By[by/IN,by/JJ,by/NN,by/RP,B-PP] Bit[bit/NN,B-NP-singular|E-NP-singular] How[how/WRB,B-ADVP] P2P[P2P/null,B-NP-singular|E-NP-singular] Is[Is/NNP,be/VBZ,B-VP] Changing[change/VBG,I-VP] the[the/DT,B-NP-singular] World[world/NN:UN,E-NP-singular] by[by/IN,B-PP] John[John/NNP,B-NP-singular] Smith[Smith/NNP,</S>,E-NP-singular]<P/>

预期提取:Bit By Bit、P2P、the World、John Smith

遇到的问题

我写的正则表达式在在线测试器能部分匹配,但运行以下Pester测试脚本时始终返回$null:

Describe "Noun Phrase Extraction" {
    $text = @"
<S> Bit[bit/NN,B-NP-singular|E-NP-singular] By[by/IN,by/JJ,by/NN,by/RP,B-PP] Bit[bit/NN,B-NP-singular|E-NP-singular] How[how/WRB,B-ADVP] P2P[P2P/null,B-NP-singular|E-NP-singular] Is[Is/NNP,be/VBZ,B-VP] Changing[change/VBG,I-VP] the[the/DT,B-NP-singular] World[world/NN:UN,E-NP-singular] by[by/IN,B-PP] John[John/NNP,B-NP-singular] Smith[Smith/NNP,</S>,E-NP-singular]<P/>
Disambiguator log:

STEP_BY_STEP[1]: Bit[bit/NN*,bite/VBD*,B-NP-singular|E-NP-singular] -> Bit[bit/NN*,B-NP-singular|E-NP-singular]

STEP_BY_STEP[1]: By[by/IN,by/JJ,by/NN,by/RP,B-PP] -> By[by/IN,by/JJ,by/NN,by/RP,B-PP]

STEP_BY_STEP[1]: Bit[bit/NN,bite/VBD,B-NP-singular|E-NP-singular] -> Bit[bit/NN,B-NP-singular|E-NP-singular]

HOW_WRB[1]: How[How/NNP,how/NN,how/WRB,B-ADVP] -> How[how/WRB,B-ADVP]

DT_JJNN_IN_NN[1]: World[world/JJ,world/NN:UN,E-NP-singular] -> World[world/NN:UN,E-NP-singular]

IN_NN_PRP[1]: by[by/IN,by/JJ,by/NN,by/RP,B-PP] -> by[by/IN,B-PP]

UPPER_NNP[1]: John[John/NNP,john/NN,B-NP-singular] -> John[John/NNP,B-NP-singular]

UPPER_NNP[1]: Smith[Smith/NNP,smith/NN,Smith/SENT_END,E-NP-singular] -> Smith[Smith/NNP,Smith/SENT_END,E-NP-singular]

5602 rules activated for language English (US)
"@

    It "should capture Noun Phrases using Regex" {
        $regexPattern = '(?<=\B)[A-Za-z]+(?=\[.*?B-NP-singular.*?\])(?:.*?)([A-Za-z]+)(?=\[.*?E-NP-singular.*?\])'

        $regexOptions = [System.Text.RegularExpressions.RegexOptions]::IgnoreCase
        $regex = New-Object System.Text.RegularExpressions.Regex $regexPattern, $regexOptions
 
        $nounPhrases = $regex.Matches($text)

        Write-Host "Matches: $($nounPhrases | Out-String)"

        foreach ($match in $nounPhrases) {
          $phrase = $match.Groups[1].Value
          Write-Host "Found phrase: $phrase" 
        }

        $capturedPhrases = @()

        foreach ($phrase in $nounPhrases) {
            $capturedPhrases += $phrase.Groups[1].Value
        }

        $capturedPhrases | Should -not -BeNullOrEmpty
    }
}

请问该正则表达式存在什么问题?


问题分析与解决方案

原正则的核心问题

  1. 边界断言错误:开头的(?<=\B)要求字符前是非单词边界,但像the这类单词前是空格(属于单词边界),直接被排除匹配,漏掉大量合法短语。
  2. 匹配范围残缺:只捕获短语的最后一个单词,没有覆盖整个短语内容;(?:.*?)无范围限制,容易跨多个无关标签导致匹配混乱。
  3. 未覆盖单单词短语:像Right-Brainers这种同时带B-NP-singular|E-NP-singular的单单词,原正则需要前后分开的B/E标签,无法匹配。
  4. 不支持带连接符的单词:[A-Za-z]+无法匹配Right-Brainers这类含-的单词,直接忽略此类短语。

修正后的正则与脚本

针对需求,用双分支正则覆盖单单词、多单词(含PP连接)的NP短语:

Describe "Noun Phrase Extraction" {
    $text = @"
<S> Bit[bit/NN,B-NP-singular|E-NP-singular] By[by/IN,by/JJ,by/NN,by/RP,B-PP] Bit[bit/NN,B-NP-singular|E-NP-singular] How[how/WRB,B-ADVP] P2P[P2P/null,B-NP-singular|E-NP-singular] Is[Is/NNP,be/VBZ,B-VP] Changing[change/VBG,I-VP] the[the/DT,B-NP-singular] World[world/NN:UN,E-NP-singular] by[by/IN,B-PP] John[John/NNP,B-NP-singular] Smith[Smith/NNP,</S>,E-NP-singular]<P/>
Disambiguator log:

STEP_BY_STEP[1]: Bit[bit/NN*,bite/VBD*,B-NP-singular|E-NP-singular] -> Bit[bit/NN*,B-NP-singular|E-NP-singular]

STEP_BY_STEP[1]: By[by/IN,by/JJ,by/NN,by/RP,B-PP] -> By[by/IN,by/JJ,by/NN,by/RP,B-PP]

STEP_BY_STEP[1]: Bit[bit/NN,bite/VBD,B-NP-singular|E-NP-singular] -> Bit[bit/NN,B-NP-singular|E-NP-singular]

HOW_WRB[1]: How[How/NNP,how/NN,how/WRB,B-ADVP] -> How[how/WRB,B-ADVP]

DT_JJNN_IN_NN[1]: World[world/JJ,world/NN:UN,E-NP-singular] -> World[world/NN:UN,E-NP-singular]

IN_NN_PRP[1]: by[by/IN,by/JJ,by/NN,by/RP,B-PP] -> by[by/IN,B-PP]

UPPER_NNP[1]: John[John/NNP,john/NN,B-NP-singular] -> John[John/NNP,B-NP-singular]

UPPER_NNP[1]: Smith[Smith/NNP,smith/NN,Smith/SENT_END,E-NP-singular] -> Smith[Smith/NNP,Smith/SENT_END,E-NP-singular]

5602 rules activated for language English (US)
"@

    It "should capture Noun Phrases using Regex" {
        # 正则覆盖两种情况:
        # 1. 单单词NP:包含B-NP-singular|E-NP-singular的单词
        # 2. 多单词NP:从B-NP-singular开始,到E-NP-singular结束,中间允许I-NP-singular或B-PP连接的结构
        $regexPattern = @'
(?:
    (\w+(?:-\w+)*)(?=\[.*?B-NP-singular\|E-NP-singular.*?\])
    |
    (\w+(?:-\w+)*)(?=\[.*?B-NP-singular.*?\])
    (?:\s+\w+(?:-\w+)*(?=\[.*?(?:I-NP-singular|B-PP).*?\]))*
    \s+(\w+(?:-\w+)*)(?=\[.*?E-NP-singular.*?\])
)
'@ -replace '\s+', '' # 去掉正则里的空白,避免影响匹配

        $regexOptions = [System.Text.RegularExpressions.RegexOptions]::IgnoreCase -bor [System.Text.RegularExpressions.RegexOptions]::Singleline
        $regex = [System.Text.RegularExpressions.Regex]::new($regexPattern, $regexOptions)
 
        $nounPhrases = $regex.Matches($text)

        $capturedPhrases = @()
        foreach ($match in $nounPhrases) {
            # 合并三个分组的内容,过滤空值
            $parts = $match.Groups | Where-Object { $_.Success -and $_.Value } | Select-Object -ExpandProperty Value
            if ($parts) {
                $capturedPhrases += $parts -join ' '
            }
        }

        Write-Host "Captured phrases: $($capturedPhrases -join ', ')"
        $capturedPhrases | Should -not -BeNullOrEmpty
        # 验证是否包含预期的短语
        $capturedPhrases | Should -Contain "Bit By Bit"
        $capturedPhrases | Should -Contain "P2P"
        $capturedPhrases | Should -Contain "the World"
        $capturedPhrases | Should -Contain "John Smith"
    }
}

修正说明

  1. 覆盖单单词短语:新增分支匹配同时带B-NP-singular|E-NP-singular的单词,用\w+(?:-\w+)*支持带连接符的单词。
  2. 正确匹配多单词短语:从B-NP-singular标记的单词开始,允许中间出现I-NP-singular(短语内修饰词)或B-PP(介词连接)的单词,直到E-NP-singular标记的单词结束。
  3. 修复边界问题:去掉错误的(?<=\B)断言,确保单词前的空格不会影响匹配。
  4. 合并分组结果:遍历匹配时拼接非空分组内容,得到完整短语。

内容的提问来源于stack exchange,提问作者quarks

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 03:35:10