You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用递归JavaScript代码解析长HTML并保留内嵌标签输出完整语句

解决HTML递归解析时保留内嵌标签的问题

你遇到的问题是因为原递归逻辑只提取了文本节点内容,没有保留元素标签的结构,导致带格式的文本被拆分。下面提供两种解决方案,既能完整保留<b>这类内嵌标签,也能适配长HTML场景。

递归实现(适合常规嵌套深度)

这个版本逻辑直观,通过递归遍历DOM节点,对文本节点和元素节点分别处理:

function getFormattedContent(node) {
  let content = '';
  for (const child of node.childNodes) {
    switch (child.nodeType) {
      // 处理文本节点,直接追加内容(可选过滤空文本)
      case Node.TEXT_NODE:
        content += child.textContent.trim() ? child.textContent : '';
        break;
      // 处理元素节点,保留标签结构并递归子节点
      case Node.ELEMENT_NODE:
        const tagName = child.tagName.toLowerCase();
        content += `<${tagName}>`;
        content += getFormattedContent(child);
        content += `</${tagName}>`;
        break;
      // 忽略注释、文档类型等非核心节点
      default:
        break;
    }
  }
  return content;
}

// 使用方法
const targetElement = document.getElementById('translate');
const fullTaggedContent = getFormattedContent(targetElement);
console.log(fullTaggedContent);

迭代实现(适配极端嵌套/超长HTML)

如果HTML嵌套深度极大,递归可能引发栈溢出,改用栈模拟递归的迭代版本更稳妥:

function getFormattedContentIterative(node) {
  let content = '';
  const stack = [{ node: node, isVisited: false }];

  while (stack.length > 0) {
    const { node: currentNode, isVisited } = stack.pop();

    if (isVisited) {
      // 二次访问时拼接结束标签
      if (currentNode.nodeType === Node.ELEMENT_NODE) {
        const tagName = currentNode.tagName.toLowerCase();
        content += `</${tagName}>`;
      }
      continue;
    }

    if (currentNode.nodeType === Node.TEXT_NODE) {
      content += currentNode.textContent.trim() ? currentNode.textContent : '';
      continue;
    }

    if (currentNode.nodeType === Node.ELEMENT_NODE) {
      // 首次访问拼接开始标签,标记待处理结束标签,子节点倒序入栈保证顺序
      const tagName = currentNode.tagName.toLowerCase();
      content += `<${tagName}>`;
      stack.push({ node: currentNode, isVisited: true });
      for (let i = currentNode.childNodes.length - 1; i >= 0; i--) {
        stack.push({ node: currentNode.childNodes[i], isVisited: false });
      }
    }
  }
  return content;
}

// 使用方法
const targetElement = document.getElementById('translate');
const fullTaggedContent = getFormattedContentIterative(targetElement);
console.log(fullTaggedContent);

额外说明

  • 如果需要保留元素的属性(如class、id),可以把元素节点的标签生成逻辑改成:content += currentNode.outerHTML.match(/^<[^>]+>/)[0];,这样能完整保留开始标签的所有属性。
  • 代码中的trim()用于过滤多余的换行和空格,不需要的话可以直接删除,保留原始文本格式。

内容的提问来源于stack exchange,提问作者corsaro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 23:20:36