使用innerText提取HTML文本时如何保留原格式?(textAngular场景)
Hey there! I get it—using innerText can be super frustrating because it strips out so much of the formatting that matters when you need to keep line breaks, spaces, and special content like math formulas intact. Let’s fix this with a targeted approach that traverses the DOM and preserves the structure you care about.
The Problem with innerText
innerText is built to return text as it would appear in a rendered browser, which means it ignores multiple spaces, collapses line breaks from block elements, and doesn’t handle special symbols consistently. Instead, we’ll use textContent (which preserves raw text including special characters) combined with manual handling of HTML tags that control formatting.
Solution: Custom DOM Traversal Function
We’ll write a recursive function that walks through each node in your parsed DOM, adds appropriate line breaks for block elements (like <p>, <div>) and <br> tags, and preserves all text content including special symbols.
Here’s how you can implement it in your Angular code:
function extractFormattedText(element) { let text = ''; const childNodes = element.childNodes; for (const node of childNodes) { if (node.nodeType === Node.TEXT_NODE) { // Preserve raw text including spaces, special symbols, and math characters text += node.textContent; } else if (node.nodeType === Node.ELEMENT_NODE) { const tagName = node.tagName.toLowerCase(); // Add line breaks for block-level elements and line breaks if (tagName === 'br') { text += '\n'; } else if (['p', 'div', 'h1', 'h2', 'h3', 'ul', 'ol'].includes(tagName)) { text += '\n'; } // Recursively process child nodes text += extractFormattedText(node); // Add an extra line break after block elements to mimic their spacing if (['p', 'div', 'h1', 'h2', 'h3', 'ul', 'ol'].includes(tagName)) { text += '\n'; } } } // Clean up excess line breaks (avoid 3+ consecutive newlines) return text.replace(/\n{3,}/g, '\n\n').trim(); } // Usage in your code const parsedDoc = parser.parseFromString(data, 'text/html'); $scope.formattedText = extractFormattedText(parsedDoc.body);
How This Works
- Text Nodes: We use
textContentinstead ofinnerTextto capture raw text, including special characters (like math symbols, Unicode characters, or escaped entities). - Element Nodes: We check for tags that affect line spacing (block elements like
<p>or<div>, or<br>tags) and add newline characters to mimic their original formatting. - Recursion: The function digs into child nodes to ensure we process every part of the HTML structure.
- Cleanup: We collapse multiple consecutive newlines to avoid messy extra spacing while keeping intentional paragraph breaks.
Adjustments for Your Needs
- If you have list items (
<li>), you can extend the function to add bullet points:if (tagName === 'li') { text += '- '; // Add bullet point for list items } - If your math formulas use MathML, this function will preserve their text content perfectly since MathML nodes are just nested elements with text that
textContentcaptures correctly.
This approach gives you full control over which formatting rules to preserve, ensuring your extracted text matches the original HTML’s structure while keeping all special symbols intact.
内容的提问来源于stack exchange,提问作者ganesh kaspate

