PHP正则表达式捕获重叠提及及上下文的实现问题
@提及检测工具的正则匹配问题解决方案
1. 可通过单次扫描实现完整匹配
可以利用零宽度环视断言实现单次扫描捕获所有提及及前后最多5个单词的上下文,避免因匹配重叠导致的遗漏。环视断言不会消耗字符串字符,正则引擎可连续扫描所有提及。
示例正则与代码:
// 匹配提及及前后最多5个单词的上下文 $pattern = '/(?<=((?:\b[\w\']+\b\s*){0,5}))(@user:\d+)(?=((?:\s*\b[\w\']+\b){0,5}))/'; // 按匹配组顺序返回结果 preg_match_all($pattern, $text, $matches, PREG_SET_ORDER); // 解析结果 $results = []; foreach ($matches as $match) { $results[] = [ 'pre_context' => trim($match[1]), 'mention' => $match[2], 'post_context' => trim($match[3]) ]; }
(?<=...)正向后顾断言:捕获提及前最多5个单词(不消耗字符)(?=...)正向前瞻断言:捕获提及后最多5个单词(不消耗字符)PREG_SET_ORDER参数让每个匹配作为独立数组返回,方便解析
2. 先定位提及位置再获取无重叠前置上下文
如果先通过preg_match_all定位所有提及,可通过控制上下文的起始范围避免包含其他提及,步骤如下:
步骤1:获取所有提及的位置与文本
// 捕获所有提及的文本和起始偏移量 preg_match_all('/@user:\d+/', $text, $matches, PREG_OFFSET_CAPTURE); $mentions = $matches[0]; $results = []; $prevMentionEnd = 0;
步骤2:遍历提取无重叠的前置上下文
foreach ($mentions as $index => $mention) { $mentionText = $mention[0]; $mentionStart = $mention[1]; $mentionEnd = $mentionStart + strlen($mentionText); // 前置上下文范围:上一个提及的结束位置 到 当前提及的起始位置 $preRaw = substr($text, $prevMentionEnd, $mentionStart - $prevMentionEnd); // 提取最多5个单词(取末尾的内容) preg_match('/(?:\b[\w\']+\b\s*){0,5}$/', trim($preRaw), $preMatch); $preContext = trim($preMatch[0] ?? ''); // 可选:提取后置上下文(当前提及结束位置到下一个提及起始位置) $nextMentionStart = isset($mentions[$index+1]) ? $mentions[$index+1][1] : strlen($text); $postRaw = substr($text, $mentionEnd, $nextMentionStart - $mentionEnd); preg_match('/^(?:\s*\b[\w\']+\b){0,5}/', trim($postRaw), $postMatch); $postContext = trim($postMatch[0] ?? ''); $results[] = [ 'mention' => $mentionText, 'pre_context' => $preContext, 'post_context' => $postContext ]; $prevMentionEnd = $mentionEnd; }
核心逻辑是:每个提及的前置上下文起始点为上一个提及的结束位置,确保上下文区间内不会包含其他提及,再从该区间内提取最多5个单词。
内容的提问来源于stack exchange,提问作者Erich
相关产品推荐
相关产品推荐

