You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则清理HTML属性时误删标签内文本的问题与实现方案

问题描述

我在解析从站点抓取的HTML内容时遇到了问题:这类HTML经常包含不规范的片段,会导致常规HTML解析器运行异常。
最初我写了一段正则,用来批量移除HTML里所有的[class]和[id]属性:

/(\[class\]((=)("|')?.*("|')))|(\[class\])|((\[id\]((=)("|')?.*("|')))|(\[id\]))/

这段正则在部分场景下可以正常工作,比如处理如下HTML时不会出错:

<div class="par fontsize-16" [class]="'par fontsize-' + fontsize"><p>the two of them left that everyone came back to their senses.</p>

但处理下面这段HTML时就会出现匹配错误:

</div><span id="saved" hidden>Settings saved..</span><div class="clear"></div><div class="par fontsize-16" [class]="'par fontsize-' + fontsize"><p>It wasn't " until the two of them left that everyone came back to their senses.</p>

错误原因是正则会把标签包裹的正文内容里的It wasn't "片段当成属性的一部分错误移除。
我的核心需求是仅删除标签内的目标属性及其对应的属性值,完全不改动标签包裹的正文文本,想确认这个需求是否可以实现。

最终解决方案

感谢@IT goldman 提供的思路,我最终调试出了可稳定运行的实现,分享出来供有同类需求的开发者参考:

function cleanHTML(html, attrs) {
  try {
    attrs.forEach(attr => {
      let pos = 0
      // 循环定位所有目标属性的位置
      while ((pos = html.indexOf(attr)) > -1) {
        let sep = null;
        let state = 0;
        // 从属性名位置向后扫描,判断属性是否带值、值的引号分隔符是什么
        for (let i = pos + attr.length; i < html.length; i++) {
          const c = html.charAt(i);
          if (c == '=') {
            state = 1
            continue;
          }
          if (state == 1 && (c.trim() === '"' || c.trim() === "'")) {
            sep = c;
            break;
          } else if (["'", '"', "=", ""].indexOf(c.trim()) === -1)
            break;
        }

        if (sep) {
          // 找到当前属性所在标签的闭合位置,避免匹配到标签外正文里的引号
          const closingPos = html.indexOf(">", pos);
          const firstQuotePos = html.indexOf(sep, pos);
          let secondQuotePos = html.indexOf(sep, firstQuotePos + 1);
          // 处理属性值引号不闭合的异常场景
          if (secondQuotePos > closingPos)
            secondQuotePos = closingPos - 1;
          html = html.substring(0, pos) + html.substring(secondQuotePos + 1)
        } else {
          // 处理无属性值的布尔属性场景
          html = html.substring(0, pos) + html.substring(pos + attr.length + (state == 1 ? 1 : 0));
        }
      }
    });
  } catch (e) {
    console.log(e);
  }
  return html;
}

// 测试用例
const testHtml = `<span [class]= [class][class] id="saved" [id]hidden [class] =  '"kjhsdf->Settings saved..</span><div class="clear"></div><div class="par fontsize-16" [class]="'par fontsize-' + fontsize"><p>It wasn't " until the two of them left that everyone came back to their senses.</p><a [class]='another'>sasportas</a>`
console.log(cleanHTML(testHtml, ["[class]", "[id]"]));

这个方案放弃了纯正则匹配的思路,改用字符串逐位扫描+边界判断的逻辑,能够准确区分标签内属性和标签外正文,不会误删正文内容,同时兼容属性值无引号、引号不闭合等多种不规范HTML场景。


内容的提问来源于stack exchange,提问作者Alen.Toma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 20:06:27