如何提取字符串中给定两个正则表达式之间的所有子串?
提取动态标记间所有子串的解决方案
核心思路
利用正则的捕获组结合全局匹配模式,将起始标记和结束标记作为动态参数,构建可匹配所有目标子串的正则表达式,再提取捕获组中的内容即可。
JavaScript 实现示例
function extractBetweenMarkers(str, startMarker, endMarker) { // 转义标记中的正则特殊字符,避免语法错误 const escapedStart = startMarker.replace(/[.*+?^${}()|[\]\\]/g, '\\$&'); const escapedEnd = endMarker.replace(/[.*+?^${}()|[\]\\]/g, '\\$&'); // 构建全局匹配正则,用非贪婪模式避免跨标记匹配 const regex = new RegExp(`${escapedStart}(.*?)${escapedEnd}`, 'g'); const matches = []; let match; // 遍历所有匹配结果,提取捕获组内容 while ((match = regex.exec(str)) !== null) { matches.push(match[1]); } return matches; } // 测试用例 console.log(extractBetweenMarkers('My $$name$$ is John, my $$surname$$ is Doe', '$$', '$$')); // 输出: ["name", "surname"] console.log(extractBetweenMarkers('My &&name&& is John, my &&surname&& is Doe', '&&', '&&')); // 输出: ["name", "surname"]
关键细节说明
- 转义特殊字符:如果标记包含
$、(、*这类正则元字符,必须先转义,否则会导致正则解析错误。 - 非贪婪匹配:使用
.*?而非.*,避免将第一个起始标记到最后一个结束标记之间的所有内容误判为单个匹配结果。 - 全局匹配:正则添加
g修饰符,才能遍历字符串中所有符合条件的子串。
Python 实现示例
import re def extract_between_markers(text, start_marker, end_marker): # 转义正则特殊字符 escaped_start = re.escape(start_marker) escaped_end = re.escape(end_marker) # 构建正则并提取所有捕获组内容 pattern = re.compile(f'{escaped_start}(.*?){escaped_end}') return pattern.findall(text) # 测试用例 print(extract_between_markers('My $$name$$ is John, my $$surname$$ is Doe', '$$', '$$')) # 输出: ['name', 'surname']
内容的提问来源于stack exchange,提问作者meursault
相关产品推荐
相关产品推荐

