如何将带嵌套标签的HTML转换为含文本与格式属性的结构化对象数组?
嵌套HTML转格式片段实现方案
核心问题
原有代码仅做了单层节点遍历,遇到嵌套标签时直接取整段textContent,没有维护逐层叠加的格式状态,因此无法处理多层格式嵌套的场景。
实现思路
- 采用深度优先遍历DOM树,维护一个格式栈存储当前遍历位置生效的所有格式标签
- 遇到行内格式标签(如
strong/em)时,先将标签名压入栈,再递归处理该标签的所有子节点,子节点处理完成后将标签弹出栈,保证格式作用域正确 - 遇到文本节点时,直接取当前格式栈的快照作为该段文本的
format值,栈为空时取null - 块级容器标签(如
p)不加入格式栈,仅作为分段容器,处理完单个块级节点后可按需追加换行标记
修正后完整代码
export interface Phrase { text: string; format: string[] | null; } // 可根据业务需要扩展格式标签白名单,不在白名单内的标签仅作为容器不记录格式 const FORMAT_TAGS = new Set(['strong', 'em', 'b', 'i', 'u', 's', 'code']); export class HTMLParser { public phrasesProcessed: Phrase[] = []; public parse(text: string): Phrase[] { this.phrasesProcessed = []; const parser = new DOMParser(); const sourceDocument = parser.parseFromString(text, "text/html"); // 初始格式栈为空 this.traverseNodes(sourceDocument.body.childNodes, []); console.log("RESULT of CONVERSION", this.phrasesProcessed); return this.phrasesProcessed; } /** * 递归遍历节点 * @param nodes 待遍历节点列表 * @param currentFormats 当前生效的格式栈 */ private traverseNodes(nodes: NodeListOf<ChildNode>, currentFormats: string[]) { Array.from(nodes).forEach((node, index) => { if (node.nodeType === Node.TEXT_NODE) { const textContent = node.textContent || ''; // 跳过完全空的文本节点,保留带空格的有效文本 if (textContent.length > 0) { this.phrasesProcessed.push({ text: textContent, format: currentFormats.length > 0 ? [...currentFormats] : null }); } } else if (node.nodeType === Node.ELEMENT_NODE && node instanceof HTMLElement) { const tagName = node.tagName.toLowerCase(); const isFormatTag = FORMAT_TAGS.has(tagName); // 是格式标签则压入栈 if (isFormatTag) { currentFormats.push(tagName); } // 递归处理子节点 this.traverseNodes(node.childNodes, currentFormats); // 处理完子节点后弹出当前格式标签 if (isFormatTag) { currentFormats.pop(); } // 原逻辑:p标签处理完后追加换行,不需要可以删掉这段 if (tagName === 'p' && index !== nodes.length - 1) { this.phrasesProcessed.push({ text: '\n', format: null }); } } }); } }
验证说明
传入示例中的HTML字符串运行后,输出结果和要求的目标结构完全一致,支持任意层级的格式标签嵌套,可通过修改FORMAT_TAGS白名单适配需要识别的格式类型。
内容的提问来源于stack exchange,提问作者thooyork
相关产品推荐
相关产品推荐

