PowerShell正则提取名词短语失败,Pester测试返回Null求助
问题描述
我正尝试用PowerShell正则表达式,从带标签和属性的特定格式字符串中提取两类名词短语:
- 由
B-NP-singular与E-NP-singular配对组成的名词短语 - 由
B-NP-singular E-NP-singular B-PP B-NP-singular E-NP-singular组合成的名词短语
输入示例与预期结果
示例1
输入:
<S> A[a/DT,B-NP-singular] Whole[whole/JJ,whole/NN,whole/RP,I-NP-singular] New[New/JJ,New/NNP,new/JJ,I-NP-singular] Mind[mind/NN:UN,E-NP-singular] Why[why/CC,why/NN,why/UH,why/WRB,B-ADVP] Right-Brainers[Right-Brainers/null,B-NP-singular|E-NP-singular] Will[will/MD,B-VP] Rule[rule/VB,I-VP] the[the/DT,B-NP-singular] Future[future/JJ,future/NN:UN,</S>,E-NP-singular]<P/>
预期提取:Whole New Mind、Right-Brainers、the Future
示例2
输入:
<S> Bit[bit/NN,B-NP-singular|E-NP-singular] By[by/IN,by/JJ,by/NN,by/RP,B-PP] Bit[bit/NN,B-NP-singular|E-NP-singular] How[how/WRB,B-ADVP] P2P[P2P/null,B-NP-singular|E-NP-singular] Is[Is/NNP,be/VBZ,B-VP] Changing[change/VBG,I-VP] the[the/DT,B-NP-singular] World[world/NN:UN,E-NP-singular] by[by/IN,B-PP] John[John/NNP,B-NP-singular] Smith[Smith/NNP,</S>,E-NP-singular]<P/>
预期提取:Bit By Bit、P2P、the World、John Smith
遇到的问题
我写的正则表达式在在线测试器能部分匹配,但运行以下Pester测试脚本时始终返回$null:
Describe "Noun Phrase Extraction" { $text = @" <S> Bit[bit/NN,B-NP-singular|E-NP-singular] By[by/IN,by/JJ,by/NN,by/RP,B-PP] Bit[bit/NN,B-NP-singular|E-NP-singular] How[how/WRB,B-ADVP] P2P[P2P/null,B-NP-singular|E-NP-singular] Is[Is/NNP,be/VBZ,B-VP] Changing[change/VBG,I-VP] the[the/DT,B-NP-singular] World[world/NN:UN,E-NP-singular] by[by/IN,B-PP] John[John/NNP,B-NP-singular] Smith[Smith/NNP,</S>,E-NP-singular]<P/> Disambiguator log: STEP_BY_STEP[1]: Bit[bit/NN*,bite/VBD*,B-NP-singular|E-NP-singular] -> Bit[bit/NN*,B-NP-singular|E-NP-singular] STEP_BY_STEP[1]: By[by/IN,by/JJ,by/NN,by/RP,B-PP] -> By[by/IN,by/JJ,by/NN,by/RP,B-PP] STEP_BY_STEP[1]: Bit[bit/NN,bite/VBD,B-NP-singular|E-NP-singular] -> Bit[bit/NN,B-NP-singular|E-NP-singular] HOW_WRB[1]: How[How/NNP,how/NN,how/WRB,B-ADVP] -> How[how/WRB,B-ADVP] DT_JJNN_IN_NN[1]: World[world/JJ,world/NN:UN,E-NP-singular] -> World[world/NN:UN,E-NP-singular] IN_NN_PRP[1]: by[by/IN,by/JJ,by/NN,by/RP,B-PP] -> by[by/IN,B-PP] UPPER_NNP[1]: John[John/NNP,john/NN,B-NP-singular] -> John[John/NNP,B-NP-singular] UPPER_NNP[1]: Smith[Smith/NNP,smith/NN,Smith/SENT_END,E-NP-singular] -> Smith[Smith/NNP,Smith/SENT_END,E-NP-singular] 5602 rules activated for language English (US) "@ It "should capture Noun Phrases using Regex" { $regexPattern = '(?<=\B)[A-Za-z]+(?=\[.*?B-NP-singular.*?\])(?:.*?)([A-Za-z]+)(?=\[.*?E-NP-singular.*?\])' $regexOptions = [System.Text.RegularExpressions.RegexOptions]::IgnoreCase $regex = New-Object System.Text.RegularExpressions.Regex $regexPattern, $regexOptions $nounPhrases = $regex.Matches($text) Write-Host "Matches: $($nounPhrases | Out-String)" foreach ($match in $nounPhrases) { $phrase = $match.Groups[1].Value Write-Host "Found phrase: $phrase" } $capturedPhrases = @() foreach ($phrase in $nounPhrases) { $capturedPhrases += $phrase.Groups[1].Value } $capturedPhrases | Should -not -BeNullOrEmpty } }
请问该正则表达式存在什么问题?
问题分析与解决方案
原正则的核心问题
- 边界断言错误:开头的
(?<=\B)要求字符前是非单词边界,但像the这类单词前是空格(属于单词边界),直接被排除匹配,漏掉大量合法短语。 - 匹配范围残缺:只捕获短语的最后一个单词,没有覆盖整个短语内容;
(?:.*?)无范围限制,容易跨多个无关标签导致匹配混乱。 - 未覆盖单单词短语:像
Right-Brainers这种同时带B-NP-singular|E-NP-singular的单单词,原正则需要前后分开的B/E标签,无法匹配。 - 不支持带连接符的单词:
[A-Za-z]+无法匹配Right-Brainers这类含-的单词,直接忽略此类短语。
修正后的正则与脚本
针对需求,用双分支正则覆盖单单词、多单词(含PP连接)的NP短语:
Describe "Noun Phrase Extraction" { $text = @" <S> Bit[bit/NN,B-NP-singular|E-NP-singular] By[by/IN,by/JJ,by/NN,by/RP,B-PP] Bit[bit/NN,B-NP-singular|E-NP-singular] How[how/WRB,B-ADVP] P2P[P2P/null,B-NP-singular|E-NP-singular] Is[Is/NNP,be/VBZ,B-VP] Changing[change/VBG,I-VP] the[the/DT,B-NP-singular] World[world/NN:UN,E-NP-singular] by[by/IN,B-PP] John[John/NNP,B-NP-singular] Smith[Smith/NNP,</S>,E-NP-singular]<P/> Disambiguator log: STEP_BY_STEP[1]: Bit[bit/NN*,bite/VBD*,B-NP-singular|E-NP-singular] -> Bit[bit/NN*,B-NP-singular|E-NP-singular] STEP_BY_STEP[1]: By[by/IN,by/JJ,by/NN,by/RP,B-PP] -> By[by/IN,by/JJ,by/NN,by/RP,B-PP] STEP_BY_STEP[1]: Bit[bit/NN,bite/VBD,B-NP-singular|E-NP-singular] -> Bit[bit/NN,B-NP-singular|E-NP-singular] HOW_WRB[1]: How[How/NNP,how/NN,how/WRB,B-ADVP] -> How[how/WRB,B-ADVP] DT_JJNN_IN_NN[1]: World[world/JJ,world/NN:UN,E-NP-singular] -> World[world/NN:UN,E-NP-singular] IN_NN_PRP[1]: by[by/IN,by/JJ,by/NN,by/RP,B-PP] -> by[by/IN,B-PP] UPPER_NNP[1]: John[John/NNP,john/NN,B-NP-singular] -> John[John/NNP,B-NP-singular] UPPER_NNP[1]: Smith[Smith/NNP,smith/NN,Smith/SENT_END,E-NP-singular] -> Smith[Smith/NNP,Smith/SENT_END,E-NP-singular] 5602 rules activated for language English (US) "@ It "should capture Noun Phrases using Regex" { # 正则覆盖两种情况: # 1. 单单词NP:包含B-NP-singular|E-NP-singular的单词 # 2. 多单词NP:从B-NP-singular开始,到E-NP-singular结束,中间允许I-NP-singular或B-PP连接的结构 $regexPattern = @' (?: (\w+(?:-\w+)*)(?=\[.*?B-NP-singular\|E-NP-singular.*?\]) | (\w+(?:-\w+)*)(?=\[.*?B-NP-singular.*?\]) (?:\s+\w+(?:-\w+)*(?=\[.*?(?:I-NP-singular|B-PP).*?\]))* \s+(\w+(?:-\w+)*)(?=\[.*?E-NP-singular.*?\]) ) '@ -replace '\s+', '' # 去掉正则里的空白,避免影响匹配 $regexOptions = [System.Text.RegularExpressions.RegexOptions]::IgnoreCase -bor [System.Text.RegularExpressions.RegexOptions]::Singleline $regex = [System.Text.RegularExpressions.Regex]::new($regexPattern, $regexOptions) $nounPhrases = $regex.Matches($text) $capturedPhrases = @() foreach ($match in $nounPhrases) { # 合并三个分组的内容,过滤空值 $parts = $match.Groups | Where-Object { $_.Success -and $_.Value } | Select-Object -ExpandProperty Value if ($parts) { $capturedPhrases += $parts -join ' ' } } Write-Host "Captured phrases: $($capturedPhrases -join ', ')" $capturedPhrases | Should -not -BeNullOrEmpty # 验证是否包含预期的短语 $capturedPhrases | Should -Contain "Bit By Bit" $capturedPhrases | Should -Contain "P2P" $capturedPhrases | Should -Contain "the World" $capturedPhrases | Should -Contain "John Smith" } }
修正说明
- 覆盖单单词短语:新增分支匹配同时带
B-NP-singular|E-NP-singular的单词,用\w+(?:-\w+)*支持带连接符的单词。 - 正确匹配多单词短语:从
B-NP-singular标记的单词开始,允许中间出现I-NP-singular(短语内修饰词)或B-PP(介词连接)的单词,直到E-NP-singular标记的单词结束。 - 修复边界问题:去掉错误的
(?<=\B)断言,确保单词前的空格不会影响匹配。 - 合并分组结果:遍历匹配时拼接非空分组内容,得到完整短语。
内容的提问来源于stack exchange,提问作者quarks
相关产品推荐
相关产品推荐

