寻求适用于AutoHotkey的两类正则表达式及代码问题解决
AutoHotkey正则匹配问题解决方案
需求说明
需要实现两类正则提取逻辑:
- 提取字符串中每个
$前后各最多7个词(支持带连字符、撇号的词,如ice-cream、ain't) - 提取每个
$前最多7个词,以及$到最近句号.之间的所有内容
原代码问题分析
- 仅匹配第一个
$:原代码中RegExMatch的起始位置使用pos + 1,未从当前匹配的结束位置开始,导致后续$无法被匹配。 - 无法提取
$后内容:原正则的词匹配逻辑(如\P{Xwd}\p{Xwd}+)存在缺陷,要求每个词前必须有非单词字符,导致连续词匹配失效。 - 弹窗内容左侧缺失字符:同样因起始位置错误(
pos + 1),每次匹配从上次匹配的第二个字符开始,截断了前面的内容。
修正后的代码实现
第一类需求:提取$前后各7个词
numm := [] test := "so this is just a test ok testing $5 dollars for that ice-cream sandwhich is not bad at all. And also another sentence to test $10 thousand for an ice-cream sand which ain't bad either." ; 正则说明:匹配$前最多7个词,$后最多7个词,支持带连字符、撇号的词 regex := "((?:\b[\p{L}\p{N}-']+\s+){0,7})\$((?:\s+[\p{L}\p{N}-']+){0,7})" pos := 0 While pos := RegExMatch(test, regex, y, pos ? pos + StrLen(y[0]) : 1) { ; 去除分组前后的多余空格 prevContent := Trim(y[1]) nextContent := Trim(y[2]) ; 组合完整匹配内容并加入数组 numm.Push(prevContent . "$" . nextContent) } ; 逐个弹窗显示匹配结果 For idx, content in numm { MsgBox, 第%idx%个匹配结果:`n%content% } return
第二类需求:提取$前7个词及$到句号间的内容
numm := [] test := "so this is just a test ok testing $5 dollars for that ice-cream sandwhich is not bad at all. And also another sentence to test $10 thousand for an ice-cream sand which ain't bad either." ; 正则说明:匹配$前最多7个词,以及$到最近句号之间的所有内容 regex := "((?:\b[\p{L}\p{N}-']+\s+){0,7})\$([^.]+)" pos := 0 While pos := RegExMatch(test, regex, y, pos ? pos + StrLen(y[0]) : 1) { prevContent := Trim(y[1]) nextContent := Trim(y[2]) numm.Push(prevContent . "$" . nextContent) } For idx, content in numm { MsgBox, 第%idx%个匹配结果:`n%content% } return
关键修正点
- 正则表达式优化:使用
\b[\p{L}\p{N}-']+准确定义词(包含字母、数字、连字符、撇号),避免原正则的匹配逻辑缺陷。 - 匹配起始位置修正:将
pos + 1改为pos + StrLen(y[0]),确保每次匹配从当前匹配的结束位置开始,覆盖所有$且不截断内容。 - 数组内容处理:对分组内容进行
Trim去除多余空格,保证输出内容格式整洁。
内容的提问来源于stack exchange,提问作者samtayl
相关产品推荐
相关产品推荐

