正则表达式需求:匹配HTML字符串中testing-options出现前的所有标题标签及文本
解决方法:匹配
testing-options宏之前的所有HTML标题 首先,你原来的正则表达式/<(h[1-6])[^>]*>(.*?)<\/\1>/gi会匹配所有H1-H6标题,但我们需要限定范围,只保留testing-options宏出现之前的那些标题。下面提供几种可行的方案:
方案1:分两步处理(简单直观)
先把宏之前的内容截取出来,再用你原有的正则匹配标题:
const htmlStr = '<h1>jfhfnjkgf</h1><h2>fegsg</h2><h2>fwegrweghrw</h2><h4>sfadfsaf</h4><h4>ukfdsnsjkfsd</h4><ac:structured-macro ac:name="testing-options" ac:schema-version="1" data-layout="default" ac:local-id="8bbfe293-eeb9-4785-9873-9fb9b218692b" ac:macro-id="bed737bbfaaf5f1df5f68a0ecc0bc6a7"><ac:parameter ac:name="enable-testing">disable</ac:parameter></ac:structured-macro><h3>bkbkvwv</h3><h4>vbdkhvbkwv</h4><h3>v n vw bv</h3><p />'; // 1. 截取到testing-options宏之前的所有内容 const contentBeforeMacro = htmlStr.split('<ac:structured-macro ac:name="testing-options"')[0]; // 2. 用原正则匹配所有标题 const titleRegex = /<(h[1-6])[^>]*>(.*?)<\/\1>/gi; const matchedTitles = contentBeforeMacro.match(titleRegex) || []; // 合并结果 const finalResult = matchedTitles.join(''); console.log(finalResult);
运行后就能得到你期望的匹配结果:<h1>jfhfnjkgf</h1><h2>fegsg</h2><h2>fwegrweghrw</h2><h4>sfadfsaf</h4><h4>ukfdsnsjkfsd</h4>
方案2:用单个正则直接匹配(一步到位)
通过正向预查来确保当前标题出现在宏之前:
const htmlStr = '<h1>jfhfnjkgf</h1><h2>fegsg</h2><h2>fwegrweghrw</h2><h4>sfadfsaf</h4><h4>ukfdsnsjkfsd</h4><ac:structured-macro ac:name="testing-options" ac:schema-version="1" data-layout="default" ac:local-id="8bbfe293-eeb9-4785-9873-9fb9b218692b" ac:macro-id="bed737bbfaaf5f1df5f68a0ecc0bc6a7"><ac:parameter ac:name="enable-testing">disable</ac:parameter></ac:structured-macro><h3>bkbkvwv</h3><h4>vbdkhvbkwv</h4><h3>v n vw bv</h3><p />'; const regex = /<h[1-6][^>]*>.*?<\/h[1-6]>(?=(?:(?!<ac:structured-macro ac:name="testing-options").)*$)/gi; const matchedTitles = htmlStr.match(regex) || []; const finalResult = matchedTitles.join(''); console.log(finalResult);
正则说明:
<h[1-6][^>]*>.*?<\/h[1-6]>:和你原正则的标题匹配逻辑一致(?=(?:(?!<ac:structured-macro ac:name="testing-options").)*$):正向预查,确保当前标题到字符串结尾的过程中,不会出现testing-options宏的起始标签,也就是标题一定在宏之前
更可靠的方案:用DOM解析(推荐)
正则处理HTML的局限性很大(比如标签换行、属性顺序变化、嵌套结构等都可能导致匹配失败),如果你的HTML结构可能更复杂,建议用DOM解析的方式:
const htmlStr = '<h1>jfhfnjkgf</h1><h2>fegsg</h2><h2>fwegrweghrw</h2><h4>sfadfsaf</h4><h4>ukfdsnsjkfsd</h4><ac:structured-macro ac:name="testing-options" ac:schema-version="1" data-layout="default" ac:local-id="8bbfe293-eeb9-4785-9873-9fb9b218692b" ac:macro-id="bed737bbfaaf5f1df5f68a0ecc0bc6a7"><ac:parameter ac:name="enable-testing">disable</ac:parameter></ac:structured-macro><h3>bkbkvwv</h3><h4>vbdkhvbkwv</h4><h3>v n vw bv</h3><p />'; // 解析HTML为DOM文档 const parser = new DOMParser(); const doc = parser.parseFromString(htmlStr, 'text/html'); // 找到testing-options宏元素 const macroNode = doc.querySelector('ac\\:structured-macro[ac\\:name="testing-options"]'); let finalResult = ''; if (macroNode) { // 收集宏之前的所有标题节点 const titleNodes = []; let currentNode = macroNode.previousSibling; while (currentNode) { if (currentNode.nodeType === Node.ELEMENT_NODE) { // 如果是标题元素,直接加入列表 if (/^H[1-6]$/i.test(currentNode.tagName)) { titleNodes.unshift(currentNode.outerHTML); } else { // 如果是其他元素,检查内部是否有标题 const nestedTitles = currentNode.querySelectorAll('h1,h2,h3,h4,h5,h6'); Array.from(nestedTitles).forEach(title => titleNodes.unshift(title.outerHTML)); } } currentNode = currentNode.previousSibling; } finalResult = titleNodes.join(''); } else { // 如果没有宏,匹配所有标题 const allTitles = doc.querySelectorAll('h1,h2,h3,h4,h5,h6'); finalResult = Array.from(allTitles).map(title => title.outerHTML).join(''); } console.log(finalResult);
这种方法能处理各种复杂的HTML结构,容错性更强。
内容的提问来源于stack exchange,提问作者Ram's
相关产品推荐
相关产品推荐

