如何用Go正则匹配含指定词及其变体的句子片段?
解决方案:提取包含指定词汇(含前缀/变体)的句子片段
问题分析
你当前的Go正则代码只能精确匹配输入的word值,无法处理以下需求:
- 匹配以
word为前缀的单词(如cont匹配continue) - 匹配
word的词形变体(如enabled匹配enable) - 正确处理句子间无空格衔接的情况(示例中
server.I这类句点直接连下一句的场景)
实现思路
- 预处理词汇:自动去除
word的常见动词后缀(d/ed/ing),得到核心词根/前缀,确保能匹配同根的所有变体。 - 构建灵活正则:
- 用
(?i)忽略大小写 - 用
\b保证单词边界,避免匹配无关子串(如cont不会匹配content) - 用
\w*匹配词根后的任意字符,覆盖前缀和变体场景 - 用
[^.!?]*匹配句子内容,支持多种结尾标点(./!/?)
- 用
- 清理结果:去除匹配片段中的多余空格,保证输出整洁
完整代码
package main import ( "fmt" "regexp" "strings" ) func extractSentences(word, sentence string) []string { cleanWord := strings.ToLower(word) // 移除常见动词后缀,提取核心部分 for _, suffix := range []string{"ing", "ed", "d"} { if strings.HasSuffix(cleanWord, suffix) { cleanWord = strings.TrimSuffix(cleanWord, suffix) break } } // 构建正则表达式,转义特殊字符避免正则注入 pattern := fmt.Sprintf(`(?i)[^.!?]*\b%s\w*\b[^.!?]*[.!?]`, regexp.QuoteMeta(cleanWord)) re := regexp.MustCompile(pattern) matches := re.FindAllString(sentence, -1) // 清理每个匹配结果的前后空格 for i := range matches { matches[i] = strings.TrimSpace(matches[i]) } return matches } func main() { sentence := "Their have a problem with the server.I have to continue the task.Hope the server will enable for everyone.Once enable then we will continue." // 测试不同输入场景 testWords := []string{"cont", "enabled", "continued", "enable"} for _, word := range testWords { fmt.Printf("输入词汇: %s\n提取的句子片段:\n%v\n\n", word, extractSentences(word, sentence)) } }
测试结果
- 输入
cont:输出["I have to continue the task.", "Once enable then we will continue."] - 输入
enabled:输出["Hope the server will enable for everyone.", "Once enable then we will continue."] - 输入
continued:输出["I have to continue the task.", "Once enable then we will continue."] - 输入
enable:输出["Hope the server will enable for everyone.", "Once enable then we will continue."]
完全符合你给出的示例需求。
内容的提问来源于stack exchange,提问作者Navjot Sharma
相关产品推荐
相关产品推荐

